Compare commits

...
Author SHA1 Message Date
JMR-dev bf2214a549 Merge remote-tracking branch 'origin/main' into merge-111-tmp 2026-08-25 20:41:26 -05:00
Jason Ross 989069207e Merge pull request #110 from JMR-dev/docs/seven-run-counts
Stop counting the run this page calls inconclusive
2026-08-25 20:32:37 -05:00
JMR-dev 72ff7adfcc Merge branch 'main' into docs/seven-run-counts 2026-08-25 20:25:06 -05:00
JMR-dev 994ea8a3dd Say in the step log that the summary was written, since nothing else can
The job summary is the deliverable #83 asked for -- "readable without opening a
log" -- and GitHub exposes no API that reads a job summary back: the check-run
output for the advisory job returns summary: null, so a write that silently did
not happen would be invisible to everything except a human on the run page. The
step log can be read, so it now carries one line saying which of the two
happened, including the case where GITHUB_STEP_SUMMARY is unset entirely, which
is what running the script by hand looks like.
2026-08-25 20:23:05 -05:00
JMR-dev 3e9528454c Announce a baseline it cannot read, rather than falling quiet
"A comparison was asked for" and "a number was found to compare against" were one
variable, and collapsing them put the report one refactor away from being the
thing #83 filed. The sed that reads FAILS_ON_EMULATOR_API37_BASELINE is anchored
at the line start, so indenting the const into an object -- or renaming it, or
moving it -- empties it, and the old code then skipped the whole comparison while
the table kept printing exactly as before. Silent, and indistinguishable from a
run that matched.

Now an unreadable baseline is itself a deviation, with the notice naming the
const so the fix is obvious. Verified against the real captured log of run
32865281555 three ways: baseline file absent, const indented into an object, and
the committed file unchanged -- the first two announce, the third stays silent.
2026-08-25 20:20:59 -05:00
JMR-dev 0702916229 Say what the advisory API 37 job actually found, so a new failure is not invisible
That job is continue-on-error and red on every PR by design, which CLAUDE.md
states plainly -- and that instruction is exactly why nobody reads it. Nothing in
a red X separates "the known three" from "the known three plus yours".

A bare failure count would not have fixed it, and this is measured rather than
assumed. The run is usually truncated: seven of eight advisory runs read on
2026-08-25 ended in `Test run failed to complete. Expected 3 tests, received 2.`
with INSTRUMENTATION_ABORTED, and one did not. A count taken from a truncated run
misleads in both directions -- a fourth marked test can still yield the same
number if the abort lands earlier, and the known set getting worse can lower it.

The test XML does not rescue it either, which was the thing worth checking before
building on it: it IS written for an aborted run, and it reports a tidy
tests="3" failures="3" for a run the runner had just described as truncated. So
the XML is the authority on how many results landed, the runner's own output is
the only authority on whether the run finished, and the report reads both and
says which number came from where.

The baseline is one number beside the marker, because the marker means "cannot
pass on this image": the count is both how many tests the advisory leg runs and
how many should fail. A smaller failure count is the interesting direction -- it
means one now passes, which is the documented trigger for deleting the
annotation.

Nothing about the job's status changes. It stays continue-on-error, stays red,
stays out of the required contexts; a deviation is a ::notice::, never an
::error::. The report is a separate script so it can be run against a real log
saved from a real CI run, which is how the comparison was shown to fire.

The gating legs get the shape without the comparison: they run the whole suite,
so comparing there would announce a deviation five times a run -- but a truncated
run reporting fewer results than it ran is what #108 looks like, and "completed
cleanly" is the field that would show it.

Closes #83
2026-08-25 16:04:13 -05:00
Jason Ross 3fd34c24a6 Merge pull request #109 from JMR-dev/test/release-permission-guard
Notice if the release job loses the permission that lets it publish
2026-08-25 15:59:04 -05:00
JMR-dev d01a46a708 Stop counting the run this page calls inconclusive
R29 found the discriminator claimed "exact across all seven" while r07 is recorded lower
down as "inconclusive rather than ruled out, because no evidence came back from it". A row
this page calls inconclusive cannot also be counted as evidence for the conclusion.

Checking it turned up a second instance of the same over-count, which R29 did not name. The
abort-cadence section said "Measured across the seven runs above" -- but the table records
r07's aborts as **not readable**, because adb wedged before a crash buffer could be taken.
Six runs contributed gaps, not seven.

Both now say six, and both say why. The discriminator paragraph also says what excluding r07
costs, which is nothing: it is a `host` row, so the discriminator predicts it would not boot,
and confirming a prediction with the one run whose evidence did not come back adds no
information in either direction. That is the point R29 made -- claiming six does not weaken
the conclusion -- and it is worth stating in the document rather than only in the ticket,
because the next reader will otherwise wonder whether a run was quietly dropped.

Deliberately left: "four of the seven runs show the directory creation itself is broken
during the loop". That is a count of how many runs showed something, not a claim that all
seven were readable for it, so it survives. Checked rather than assumed, and named here so
the next pass does not re-audit it.

R29's other half -- "state how r07's boot outcome was read" -- is not taken, because I do
not know and inventing a source would be worse than narrowing the claim. Narrowing is the
option R29 offered and the one that can be honest.

Closes #38.
2026-08-25 15:54:48 -05:00
JMR-dev 3806641cb2 Merge branch 'main' into test/release-permission-guard 2026-08-25 15:50:20 -05:00
Jason Ross 5f9498150c Merge pull request #106 from JMR-dev/ci/build-workflow-permissions
Declare build.yml's token reach in build.yml
2026-08-25 15:49:19 -05:00
JMR-dev 8d8703ab49 Merge branch 'main' into ci/build-workflow-permissions 2026-08-25 15:41:39 -05:00
Jason Ross 27b7654418 Merge pull request #105 from JMR-dev/docs/api37-point-release
Say 37.0 is the choice, not the only api-level that exists
2026-08-25 15:41:11 -05:00
JMR-dev 4a8e30099e Notice if the release job loses the permission that lets it publish
build.yml's `release` job declares `contents: write`, and nothing checked it. Deleting those
lines leaves actionlint clean and CodeQL silent -- a narrower permission is not an alert --
and the job is `if: startsWith(github.ref, 'refs/tags/v')`, so no pull request and no merge
can exercise it. Measured with the declaration removed: every gating check still passed. The
first thing that would notice is a release failing to publish, at the moment someone is
trying to cut one.

The deletion also looks like tidying. #106 has just put a top-level `permissions: contents:
read` directly above it, so a reader could reasonably take the job-level block for a
duplicate. It is an override, and a comment saying so is not a check.

BackupExclusionsTest is the precedent: configuration rather than code, load bearing, and
unguarded because nothing compiles it.

The part worth reading twice is the second commit-worth of work in here. The test passed,
and then the mutation that is supposed to redden it did not:

  BUILD SUCCESSFUL in 614ms

Gradle cannot infer that a test depends on a file outside the source set, so the task stayed
UP-TO-DATE and the test never ran. Under --rerun-tasks the same mutation failed it properly,
which is the tell: the assertion was right and the wiring was not. A guard that does not
re-run when its subject changes is not a guard -- it is a test that will be green on the day
it matters, which is worse than no test because it reads as cover.

Fixed by declaring the workflow as a task input. Verified the whole way round afterwards,
without --rerun-tasks: mutate the file and the task re-runs and fails; restore it and the
task re-runs and passes.

What this pins and what it does not: it asserts the declaration exists in the release job's
block. It cannot assert a release actually publishes -- that needs a tag push, which is the
thing no PR can do. A tripwire against silent removal, not proof the path works, and the
KDoc says so.

Closes #107.
2026-08-25 15:38:14 -05:00
JMR-dev e7d84cc69f Merge branch 'main' into docs/api37-point-release 2026-08-25 15:33:27 -05:00
Jason Ross a8b494b846 Merge pull request #104 from JMR-dev/docs/benchmark-populate-path
Stop telling people to stage the benchmark the one way it cannot be staged
2026-08-25 15:33:02 -05:00
JMR-dev 865a4a7c8e Merge branch 'main' into docs/benchmark-populate-path 2026-08-25 10:21:23 -05:00
Jason Ross e856679395 Merge pull request #103 from JMR-dev/fix/dead-assertion-probe-test
Delete an assertion that could never fail, and say what guards instead
2026-08-25 10:21:18 -05:00
JMR-dev 49c483d877 Declare build.yml's token reach in build.yml
CodeQL alert #1, the only open one on this repository:

  actions/missing-workflow-permissions, warning / medium, build.yml:23
  Actions job or workflow does not limit the permissions of the GITHUB_TOKEN.

Alerts 2, 3 and 4 were the same rule against status_check.yml and are fixed -- that file
has a top-level block. build.yml declares permissions in exactly one place, the release
job's `contents: write`, and has no top-level default, so the `test` job inherits the
repository setting.

**Nothing is over-privileged today.** The repository default is already `read`
(default_workflow_permissions: read, can_approve_pull_request_reviews: false, read from the
API rather than assumed), so the test job holds a read token now. Saying so matters: this
is hygiene, and a commit that implied it was closing a live hole would be overstating it.

What it buys is that the default CANNOT widen these jobs later without someone editing this
file. That is not invented for the occasion -- it is the argument status_check.yml already
makes, which even names this file:

  the token's reach should be readable here, and a default that widens later should not
  silently widen these jobs with it. build.yml's release job makes the opposite
  declaration for the same reason.

So the principle was decided, applied in two workflows and in one job of this one, and the
top level of build.yml was the gap.

Verified the thing that would actually break: the release job's `contents: write` still
wins. Top level is a default, not a ceiling -- parsed and printed both, test inherits
`contents: read`, release keeps `contents: write`.

Also ran the ticket's mutation, and it found something. Deleting the release job's
`contents: write` leaves actionlint green and CodeQL quiet -- a narrower permission is not
an alert -- so nothing would catch it until a tagged release failed to publish. That is a
separate gap and is filed rather than fixed here.

actionlint clean at the pinned digest. Comment and permissions only; no step, job or
trigger changes.

Closes #100.
2026-08-25 10:16:08 -05:00
JMR-dev 1b220856ab Say 37.0 is the choice, not the only api-level that exists
R19 raised two things about this comment. One resolved itself: it used to explain why the
matrix had no API 37 row at all, and #56 added the gating row, so that half is gone.

The other survived, and this is it. The comment read

  api-level must be "37.0". A bare 37 is not an SDK package and fails during setup

The second sentence is true and was measured -- it cost a run to find. The first overstates
it. What must be true is that the api-level is a POINT release; 37.0 is one of several.
api37-debug.yml's own input descriptions already say so:

  API level, as the SDK spells it. 37.0, 37.1, 37.2-beta3, 36 ...
  System image target. android-37.1 and 37.2-beta* ship ONLY as google_apis_ps16k

and docs/api-37-emulator-crash.md measures android-37.0 rev 6 and android-37.1 rev 8 side
by side, both aborting. So the repo already knows 37.1 exists and behaves the same; only
this comment implied otherwise.

That matters for the reader it is written for. Someone debugging this row and wondering
whether a newer image helps reads "must be 37.0" as a constraint and stops. The measured
answer is that it does not help, which is a better thing to learn than a rule that is not
one -- and the ps16k-only wrinkle above 37.0 is the detail that would actually bite them.

Comment only. No job, matrix, filter or gating behaviour changes. actionlint clean at the
pinned digest.

Closes #28.
2026-08-25 10:14:34 -05:00
JMR-dev 40ae524388 Merge branch 'main' into fix/dead-assertion-probe-test 2026-08-25 10:13:12 -05:00
Jason Ross 58a29ab093 Merge pull request #99 from JMR-dev/ci/actionlint
Lint the bash inside the workflows, not only the bash in files
2026-08-25 10:13:00 -05:00
JMR-dev a1d79c212a Merge branch 'main' into ci/actionlint 2026-08-25 09:51:15 -05:00
Jason Ross bc66906dc3 Merge pull request #98 from JMR-dev/test/device-codecs-encode-consequence
Hold the two codec MIME claims that only existed in prose
2026-08-25 09:51:00 -05:00
JMR-dev d37c391c60 Stop telling people to stage the benchmark the one way it cannot be staged
RealMediaBenchmark's class KDoc said:

  Populate with:
    adb push <file>.mp4 /sdcard/Android/data/org.libremediaconverter/files/

Twelve lines below, the `samples` property KDoc -- on `get() = context.filesDir` -- says:

  Internal storage, not the external files dir. Files placed in the external dir by
  `adb push` or `adb shell cp` stay owned by the shell user, and the app then gets
  EACCES trying to read them -- which presents as an unparseable input rather than a
  permission problem.

Different directories, and the second exists specifically to explain why the first fails.
Anyone following the class KDoc stages files the benchmark cannot read, gets a skip, and
reads the skip as "not staged yet" -- the failure mode the property KDoc warns about, walked
into by the instruction in the same file.

The fix is not a corrected command. Restating the mechanism in a second place is what let
these drift, and a replacement command I have not executed would be the same defect with a
fresher date. The class KDoc now names [samples] as the single place that answers it.

Two things added that are checkable rather than remembered: the exact filenames the tests
look for, via [H264_SAMPLE] and [AV1_SAMPLE] -- the old text said `<file>.mp4`, so even the
right directory left you guessing -- and a note that the two skips every green E2E leg
reports are these.

Not claimed: that the benchmark misbehaves on CI. An earlier version of the ticket said so;
it was wrong, and measuring settled it -- both tests report SKIPPED on the gating legs, the
guards work, and "harmless in CI" is accurate. The failure that prompted the look is
Media3EngineTest, tracked as #102.

Closes #101.
2026-08-25 09:45:00 -05:00
JMR-dev 95902a7889 Delete an assertion that could never fail, and say what guards instead
ConversionViewModelProbeFailureTest's pickedProbe() helper held:

  val ready = awaitState(viewModel.state, "Ready with a probe") {
      it is ConversionState.Ready && it.input.probe != null
  }
  assertNull("nothing here should reach a terminal failure", (ready as? ConversionState.Failed))

The predicate requires `Ready`. `Ready` and `Failed` are sibling subtypes of one sealed
interface, so `ready as? Failed` is always null and the assertNull could never fire. R26
filed this PLAUSIBLE on types read; it is measured now.

Flipping the line to assertNotNull failed 3 of the 4 tests in the class -- three, because
pickedProbe() has three callers, which is also why a dead line here was worth removing
rather than shrugging at: it read as coverage in a helper the whole class depends on.

Deleted rather than replaced. There is nothing for a live assertion to add: a pick that
ended in Failed never satisfies the predicate, so awaitState fails on its timeout naming
what it was waiting for -- "Ready with a probe" -- which is a better failure message than
the assertion would have produced. The comment now says that, so the next reader does not
re-add the guard the predicate already is.

This is the ninth vacuous assertion this line of work has turned up, and the pattern is
consistent: they hide in helpers, they pass, and they look like care. The suite is green
before and after, which is exactly the point -- deleting a dead assertion cannot change a
result, and if it had, the line was not dead.

Closes #35.
2026-08-25 09:36:36 -05:00
JMR-dev 3f140fc2b1 Lint the bash inside the workflows, not only the bash in files
The shellcheck step added a few hours ago reads `git ls-files '*.sh'`. That is four files.
It does not read the inline `run:` blocks, and a good deal of this repo's bash lives there:
the release verification in build.yml, the emulator setup and teardown in status_check.yml
and api37-debug.yml. "shellcheck runs in CI" was true of the files and not of the blocks,
and CLAUDE.md said so rather than pretending otherwise.

actionlint closes that half. It parses each workflow and runs shellcheck over every `run:`,
on top of its own checks for expression syntax, `needs:` references, matrix keys and action
input names.

Pinned by digest, for the reason shellcheck is pinned -- a new rule making untouched files
fail is a red build whose diff cannot explain it -- and for a second reason of its own.
actionlint's documented install is

  bash <(curl -s https://raw.githubusercontent.com/.../download-actionlint.bash)

off a moving branch. Running that in a repository that pins every action by SHA would
contradict its own supply-chain posture more than the linter is worth. That is why #70 was
filed instead of bolted onto the shellcheck commit.

It reported exactly one finding, and it is fixed here rather than suppressed: build.yml
parsed `ls` to pick the release APK (SC2012). The glob was already in the line, so a bash
array reads it without the pipe. Gradle's output names have no spaces today, which is the
kind of assumption that holds right up until it does not.

Proved it catches something, rather than trusting a green run: planting `if [ $UNQUOTED =
bad ]` into a build.yml `run:` block produces

  shellcheck reported issue in this script: SC2086:info:4:6:

Removed again afterwards. A linter that cannot be shown to catch a plant is not wired in,
it is just running -- and SC2086 in a `run:` block is invisible to the .sh-file step, which
is the whole argument for this commit.

CLAUDE.md loses the "does not cover inline run: blocks" caveat, because it no longer does.
Both linters verified clean at their pinned digests.

Closes #70.
2026-08-25 00:26:07 -05:00
11 changed files with 532 additions and 16 deletions
+294
View File
@@ -0,0 +1,294 @@
#!/usr/bin/env bash
#
# Reports the SHAPE of an instrumented run -- how many tests were expected, how many
# reported, how many failed, and whether the run completed at all -- to the step log and to
# the job summary. In advisory mode it also compares that shape against a committed baseline
# and says plainly whether it matches.
#
# WHY THIS EXISTS (#83): the advisory API 37 leg is red on every PR by design, so a NEW failure
# joining the known ones is invisible -- nothing in a red X distinguishes "the known ones" from
# "the known ones plus yours". CLAUDE.md tells everyone not to read that job's red as their
# change breaking something, which is correct, and which also means nobody looks.
#
# WHY NOT A BARE FAILURE COUNT, measured rather than assumed. On this image the run is usually
# truncated: `Test run failed to complete. Expected 3 tests, received 2.` with
# `INSTRUMENTATION_ABORTED: System has crashed.` A count taken from a truncated run misleads in
# both directions -- a fourth marked test can still yield the same number if the abort lands
# earlier, and the known set getting worse can LOWER it. So all four fields are recorded, and
# the one saying the run was truncated is recorded with them.
#
# WHY IT IS A SEPARATE SCRIPT rather than a function inside e2e-run.sh: it is a pure seam. It
# reads a captured log plus the test XML and writes a report, so it can be run against a REAL
# log saved from a REAL CI run -- which is how the baseline comparison was shown to fire
# without waiting on an emulator. `git ls-files '*.sh'` also picks it up for shellcheck for
# free.
#
# THIS SCRIPT NEVER FAILS A RUN. It is a diagnostic, and e2e-run.sh's header explains why that
# rule is absolute here. Every field defaults to `unknown` and every comparison is guarded,
# because an unset variable under `set -u`, or a `[ "" -eq 3 ]`, is exactly how a diagnostic
# becomes the thing that turns a leg red. It exits 0 unconditionally.
#
# Usage:
# e2e-report-shape.sh <label> <gradle-log> [<baseline-file>]
#
# With a third argument the run is compared against the baseline in that file (advisory mode)
# and a `::notice::` is emitted per deviation. NEVER `::error::`: the advisory job is
# `continue-on-error: true` and stays that way, and an error annotation would be a new way for
# a diagnostic to change a conclusion.
set -uo pipefail
LABEL="${1:-unknown}"
LOG="${2:-}"
BASELINE_FILE="${3:-}"
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
REPO_ROOT="$(cd -- "$SCRIPT_DIR/../.." && pwd)"
XML_DIR="$REPO_ROOT/app/build/outputs/androidTest-results/connected/debug"
# Gradle colours its output even when it is piped, so `FAILED` arrives wrapped in escape codes.
# The numeric lines parsed below are not coloured, but stripping is cheap insurance against a
# pattern that would otherwise silently match nothing.
ESC="$(printf '\033')"
scan() { [ -s "$LOG" ] && sed -e "s/${ESC}\[[0-9;]*[a-zA-Z]//g" -- "$LOG"; }
first_number() { grep -oE '[0-9]+' | head -1; }
# ---------------------------------------------------------------------------
# Source 1: the runner's own output. This is the ONLY place a truncation is visible. The test
# XML read below is written even for an aborted run and says nothing whatever about the abort
# -- measured on run 32865281555, where the XML reports a tidy tests="3" failures="3" for a run
# the runner had just described as truncated. That is the reason this parses stdout at all.
# ---------------------------------------------------------------------------
starting_line="$(scan | grep -aoE 'Starting [0-9]+ tests on .*' | tail -1)"
abort_line="$(scan | grep -aoE 'Test run failed to complete\. Expected [0-9]+ tests, received [0-9]+\.' | tail -1)"
aborted_hits="$(scan | grep -ac 'INSTRUMENTATION_ABORTED' || true)"
failure_line="$(scan | grep -aoE 'There was [0-9]+ failure\(s\)\.' | tail -1)"
failed_names="$(scan | grep -aoE 'Execute [A-Za-z0-9_.$]+: FAILED' | sed -e 's/^Execute //' -e 's/: FAILED$//' | sort -u)"
expected="$(printf '%s' "$starting_line" | first_number)"
expected_src="\`$starting_line\`"
abort_expected="$(printf '%s' "$abort_line" | grep -oE 'Expected [0-9]+' | first_number)"
abort_received="$(printf '%s' "$abort_line" | grep -oE 'received [0-9]+' | first_number)"
log_failed="$(printf '%s' "$failure_line" | first_number)"
# `Starting N tests` is missing when the framework restarted under the run and Gradle never got
# a test list. The truncation line still carries the number it was told to expect.
if [ -z "$expected" ] && [ -n "$abort_expected" ]; then
expected="$abort_expected"
expected_src="\`$abort_line\`"
fi
# ---------------------------------------------------------------------------
# Source 2: the JUnit XML. Measured on both a truncated advisory run and a green gating leg:
# `<testsuites tests="N" failures="M">` is present in both, and aggregates every suite. It is
# the authority on how many results landed and how many were failures. It is NOT an authority
# on whether the run finished, which is what source 1 is for.
# ---------------------------------------------------------------------------
#
# Read ONLY when the runner said a test run happened. `app/build` survives between runs on a
# developer machine -- tools/local-emulator/run-e2e.sh drives several API levels against one
# checkout -- so a leg that never got as far as starting tests would otherwise be reported from
# the previous leg's XML, which is a wrong answer rather than a missing one.
xml_head=""
xml_count=0
if [ -n "$starting_line$abort_line" ] && [ -d "$XML_DIR" ]; then
while IFS= read -r f; do
xml_count=$((xml_count + 1))
[ -z "$xml_head" ] && xml_head="$(grep -ao '<testsuites[^>]*>' "$f" | head -1)"
done < <(find "$XML_DIR" -maxdepth 1 -name 'TEST-*.xml' -print 2> /dev/null | sort)
fi
xml_tests="$(printf '%s' "$xml_head" | grep -oE ' tests="[0-9]+"' | first_number)"
xml_failed="$(printf '%s' "$xml_head" | grep -oE ' failures="[0-9]+"' | first_number)"
# ---------------------------------------------------------------------------
# Derive the four fields, each with where its number came from. Everything stays a string, so a
# missing source reads `unknown` rather than becoming 0 -- a report claiming 0 tests when it
# merely could not see them would announce a deviation on every cancelled run.
# ---------------------------------------------------------------------------
received="unknown"
received_src="no source"
if [ -n "$xml_tests" ]; then
received="$xml_tests"
received_src="test XML \`<testsuites tests=\"$xml_tests\">\`"
elif [ -n "$abort_received" ]; then
received="$abort_received"
received_src="\`$abort_line\`"
elif [ -n "$expected" ] && [ -z "$abort_line" ]; then
received="$expected"
received_src="the run was not truncated, so every expected test reported"
fi
failed="unknown"
failed_src="no source"
if [ -n "$xml_failed" ]; then
failed="$xml_failed"
failed_src="test XML \`<testsuites failures=\"$xml_failed\">\`"
elif [ -n "$log_failed" ]; then
failed="$log_failed"
failed_src="\`$failure_line\`"
fi
if [ -z "$expected" ]; then
expected="unknown"
expected_src="no \`Starting N tests\` line"
fi
# A run whose start nobody can see is not a run of zero tests. Cancellation (this workflow sets
# cancel-in-progress) and the `Starting 0 tests` shape a framework restart produces both land
# here, and both have to say so rather than compare a number that does not exist.
no_run="none"
if [ "$expected" = "unknown" ] && [ "$received" = "unknown" ]; then
no_run="nothing"
elif [ "$expected" = "0" ]; then
no_run="zero"
fi
if [ -n "$abort_line" ]; then
completed="**no**"
completed_src="\`$abort_line\` with \`INSTRUMENTATION_ABORTED\`"
elif [ "${aborted_hits:-0}" -gt 0 ]; then
completed="**no**"
completed_src="\`INSTRUMENTATION_ABORTED\` in the runner output"
elif [ "$no_run" = "nothing" ]; then
# "cleanly" would be a lie about a run that left no evidence it happened.
completed="unknown"
completed_src="no runner output to read"
else
completed="yes"
completed_src="no truncation line and no \`INSTRUMENTATION_ABORTED\`"
fi
# ---------------------------------------------------------------------------
# Advisory mode: compare against the committed baseline.
#
# ONE number covers both compared fields, and that is deliberate rather than a shortcut. The
# marker means "cannot pass on this image", so the number of tests carrying it is both how many
# the advisory leg should run and how many should fail. A smaller `failed` means one now passes
# -- which is the trigger to delete the annotation, written down in FailsOnEmulatorApi37.kt.
# ---------------------------------------------------------------------------
#
# `advisory` and `baseline` are two variables on purpose. "A comparison was asked for" and "a
# number was found to compare against" are different facts, and collapsing them is how this
# report would go quietly back to being the thing #83 filed: the `sed` below is anchored, so
# indenting the const into an object -- or renaming it, or moving it to another file -- empties
# `baseline`, and a single flag would take the whole comparison down with it while the table
# kept printing. An unreadable baseline is itself a deviation, and is announced as one.
advisory="no"
baseline=""
marked=""
deviations=()
if [ -n "$BASELINE_FILE" ]; then
advisory="yes"
[ -f "$BASELINE_FILE" ] \
&& baseline="$(sed -nE 's/^const val FAILS_ON_EMULATOR_API37_BASELINE = ([0-9]+).*/\1/p' "$BASELINE_FILE" | head -1)"
# The #81 check, verbatim: what the tree actually carries. Reported next to the baseline so a
# stale baseline shows up here rather than only once the emulator disagrees with it.
if [ -d "$REPO_ROOT/app/src/androidTest" ]; then
marked="$(grep -rn "@FailsOnEmulatorApi37" "$REPO_ROOT/app/src/androidTest" --include='*.kt' \
| grep -v import | grep -c FailsOn || true)"
fi
fi
if [ "$advisory" = "yes" ] && [ -z "$baseline" ]; then
deviations+=("the committed baseline could not be read from \`$(basename -- "$BASELINE_FILE")\` — has \`FAILS_ON_EMULATOR_API37_BASELINE\` been renamed, indented into a class, or moved? Nothing was compared")
fi
if [ -n "$baseline" ]; then
if [ "$no_run" = "nothing" ]; then
deviations+=("no test run observed — the runner never reported starting one, where the baseline expects $baseline tests carrying \`@FailsOnEmulatorApi37\`")
elif [ "$no_run" = "zero" ]; then
deviations+=("the runner started 0 tests, where the baseline expects $baseline — on this image that is the framework having restarted under the run, not an empty test list")
else
if [ "$expected" != "unknown" ] && [ "$expected" != "$baseline" ]; then
deviations+=("the runner started $expected tests, the baseline is $baseline")
fi
if [ "$failed" != "unknown" ] && [ "$failed" != "$baseline" ]; then
deviations+=("$failed tests failed, the baseline is $baseline — every test carrying the marker is expected to fail on this image, so fewer means one now passes and more means a new one joined")
fi
fi
if [ -n "$marked" ] && [ "$marked" != "$baseline" ]; then
deviations+=("the tree carries $marked tests marked \`@FailsOnEmulatorApi37\` but the baseline says $baseline — update FAILS_ON_EMULATOR_API37_BASELINE")
fi
fi
# ---------------------------------------------------------------------------
# Emit. Step log first, so the common case needs neither the summary page nor an artifact.
# ---------------------------------------------------------------------------
echo "----- RUN SHAPE (api${LABEL}) -----"
echo " expected: $expected"
echo " received: $received"
echo " failed: $failed"
echo " completed cleanly: ${completed//\*/}"
if [ -n "$abort_received" ]; then
echo " received before the abort: $abort_received"
fi
if [ -n "$failed_names" ]; then
echo " failed tests:"
printf '%s\n' "$failed_names" | sed -e 's/^/ /'
fi
if [ "$advisory" = "yes" ]; then
if [ "${#deviations[@]}" -eq 0 ]; then
echo " baseline: matches ($baseline expected, $baseline failed)"
else
printf ' baseline DEVIATION: %s\n' "${deviations[@]}"
fi
fi
# A notice, never an error. See the header.
if [ "${#deviations[@]}" -gt 0 ]; then
for d in "${deviations[@]}"; do
echo "::notice::E2E api${LABEL}: $d"
done
fi
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
{
echo "### E2E api${LABEL} — run shape"
echo
echo "| field | value | where it came from |"
echo "| --- | --- | --- |"
echo "| expected | $expected | $expected_src |"
echo "| received | $received | $received_src |"
echo "| failed | $failed | $failed_src |"
echo "| completed cleanly | $completed | $completed_src |"
if [ -n "$abort_received" ]; then
echo "| received before the abort | $abort_received | the same line — the XML above counts the truncated test as a failure, this number does not |"
fi
echo
if [ -n "$failed_names" ]; then
echo "Failed:"
echo
printf '%s\n' "$failed_names" | sed -e 's/^/- `/' -e 's/$/`/'
echo
fi
if [ "$xml_count" -gt 1 ]; then
echo "> $xml_count test XML files were present; the counts above come from the first."
echo
fi
if [ "$advisory" = "yes" ]; then
if [ "${#deviations[@]}" -eq 0 ]; then
echo "**Matches the committed baseline of $baseline** — $baseline tests carry \`@FailsOnEmulatorApi37\` and all $baseline failed, which is what this job is for."
elif [ -z "$baseline" ]; then
echo "**The committed baseline could not be read, so nothing was compared.** Announced as a notice, not an error: this job is advisory and its conclusion is unchanged by anything here."
echo
printf -- '- %s\n' "${deviations[@]}"
else
echo "**DEVIATION from the committed baseline of $baseline.** Announced as a notice, not an error: this job is advisory and its conclusion is unchanged by anything here."
echo
printf -- '- %s\n' "${deviations[@]}"
fi
echo
echo "<sub>The baseline lives beside the marker, in \`FailsOnEmulatorApi37.kt\`. \`completed cleanly\` is recorded rather than compared: the truncation is intermittent — of eight advisory runs read on 2026-08-25, seven aborted and one did not — so comparing it would announce a deviation on a run that is fine.</sub>"
else
echo "<sub>No baseline comparison: that is the advisory API 37 leg only. The shape is recorded here anyway because a truncated run reports fewer results than it ran, which is what issue #108 looks like on a gating leg.</sub>"
fi
echo
} >> "$GITHUB_STEP_SUMMARY"
# The summary page is the deliverable -- "readable without opening a log" is what #83 asked
# for -- and GitHub exposes no API for reading a job summary back, so a write that silently
# did not happen would be invisible. This line is in the step log, which can be read.
echo " (the table above is also on the job summary page)"
else
echo " (GITHUB_STEP_SUMMARY is unset -- step log only)"
fi
exit 0
+36 -1
View File
@@ -34,6 +34,13 @@ TMP="${RUNNER_TEMP:-/tmp}"
LOGCAT_LOG="$TMP/logcat-api${LABEL}.txt"
DIAG_LOG="$TMP/diagnostics-api${LABEL}.txt"
WEDGE_LOG="$TMP/wedge-diagnostics-api${LABEL}.txt"
# Gradle's own output, captured to a file as well as the step log, because the run-shape report
# below has to parse it. Uploaded with the diagnostics, so a report that reads wrong can be
# checked against what it read.
GRADLE_LOG="$TMP/gradle-api${LABEL}.txt"
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
REPO_ROOT="$(cd -- "$SCRIPT_DIR/../.." && pwd)"
# ~5 min is a healthy leg (measured across API 33-36), and this wraps only the gradle client,
# a subset of that. 20 min is generous enough never to trip on a slow-but-working run, and far
@@ -214,12 +221,40 @@ status=0
# script rather than forking it: that runs several API levels back to back against one checkout
# and passes `--rerun`, so a level cannot be skipped as up-to-date and report the previous
# level's results as its own. CI gets a fresh runner per level and does not need it.
#
# `2>&1 | tee`, and the `2>&1` is the load-bearing half. The step log merges both streams, so
# reading one cannot tell you which stream a line came from -- and the single line the report
# below needs most, `Test run failed to complete. ... INSTRUMENTATION_ABORTED`, is not on
# stdout. Capturing stdout alone would leave the report saying "completed cleanly: yes" forever,
# which is precisely the comparison that cannot fire. pipefail is already set and tee exits 0,
# so the pipeline's status is still gradle's -- including the 124 that means the wrapper fired.
#
# `tee` and not `tee -a`, unlike the logcat above: CI gets a fresh runner per leg, but
# tools/local-emulator/run-e2e.sh reuses one machine, and an appended log would have the report
# reading the PREVIOUS run of the same API level. The console goes plain rather than showing
# gradle's live progress bar, which is what it already did in CI.
# shellcheck disable=SC2086
timeout -k 30s "$WEDGE_TIMEOUT" \
./gradlew :app:connectedDebugAndroidTest -PabiFilters=x86_64 --stacktrace \
${E2E_EXTRA_GRADLE_ARGS:-} || status=$?
${E2E_EXTRA_GRADLE_ARGS:-} 2>&1 | tee "$GRADLE_LOG" || status=$?
echo "::endgroup::"
# The run-shape report: expected/received/failed and whether the run finished, every time,
# green or red. It never changes `status` -- it is a diagnostic, and the header's rule about
# diagnostics applies to it as much as to every probe below.
#
# The baseline argument, and only it, turns on the comparison, and only the advisory API 37 job
# passes E2E_ADVISORY=1. Comparing on the gating legs would announce a deviation on all five of
# them every run, since they run the whole suite rather than the marked three. They still get
# the report: a truncated run reporting fewer results than it ran is what #108 looks like, and
# `completed cleanly` is the field that shows it.
if [ "${E2E_ADVISORY:-}" = "1" ]; then
bash "$SCRIPT_DIR/e2e-report-shape.sh" "$LABEL" "$GRADLE_LOG" \
"$REPO_ROOT/app/src/androidTest/java/org/libremediaconverter/FailsOnEmulatorApi37.kt" || true
else
bash "$SCRIPT_DIR/e2e-report-shape.sh" "$LABEL" "$GRADLE_LOG" || true
fi
if [ "$status" -eq 0 ]; then
kill "$LOGCAT_PID" 2>/dev/null || true
exit 0
+16 -1
View File
@@ -13,6 +13,17 @@ on:
# reference amounts to running whatever that repository contains tomorrow. This matters
# more here than on pull requests: these jobs sign nothing today, but they do publish
# the artifacts people install.
# Declared here rather than inherited, for the reason status_check.yml gives for its own
# block: the token's reach should be readable in the file that uses it, and a repository
# default that widens later should not silently widen these jobs with it. The repository
# default is `read` today, so this changes nothing about what runs -- it fixes what a
# reader can know without leaving the file, and it is what CodeQL alert #1 asked for.
#
# The `release` job below overrides this with `contents: write`, which is how job-level
# permissions work: this is a default, not a ceiling.
permissions:
contents: read
env:
GRADLE_CACHE_PATHS: |
~/.gradle/caches
@@ -81,7 +92,11 @@ jobs:
- name: Verify the released artifacts
run: |
APK=$(ls app/build/outputs/apk/release/*.apk | head -1)
# A glob, not `ls | head`: the glob is already here, and parsing ls is what
# SC2012 is about. Gradle's names have no spaces today, which is exactly the
# kind of assumption that holds until it does not.
apks=(app/build/outputs/apk/release/*.apk)
APK="${apks[0]}"
# A release that shipped one ABI, or lost 16 KB alignment, would install
# fine on a test device and fail for users or at Play submission. Both are
# cheap to check and expensive to discover later.
+32 -2
View File
@@ -182,6 +182,23 @@ jobs:
docker run --rm "$SHELLCHECK" --version
git ls-files -z '*.sh' | xargs -0 -r docker run --rm -v "$PWD:/mnt" "$SHELLCHECK"
# actionlint closes the half shellcheck cannot see. The step above reads .sh files;
# a good deal of this repo's bash lives in inline `run:` blocks instead -- the release
# verification here, the emulator setup and teardown in this file and in
# api37-debug.yml. actionlint parses each workflow and runs shellcheck over every
# `run:`, on top of its own checks for expression syntax, `needs:` references, matrix
# keys and action input names.
#
# Pinned by digest for the same reason shellcheck is, and with a second reason of its
# own: actionlint's documented install is `bash <(curl -s .../download-actionlint.bash)`
# off a moving branch, which would sit badly in a repo that pins every action by SHA.
- name: actionlint
env:
ACTIONLINT: rhysd/actionlint@sha256:9d36088643581e728c969f35141f88139fec77280b2be23c1f66f8e40e1025e7
run: |
docker run --rm "$ACTIONLINT" -version
docker run --rm -v "$PWD:/repo" -w /repo "$ACTIONLINT" -color
# `!cancelled()` rather than a plain sequence: a shellcheck failure above must not
# cost the ktlint/detekt/lint lists. Same reason this step passes --continue -- one
# round trip should produce every list, not stop at the first.
@@ -263,8 +280,13 @@ jobs:
# docs/api-37-emulator-crash.md has the per-method measurements, and the
# correction that produced them.
#
# api-level must be "37.0". A bare 37 is not an SDK package and fails
# during setup, which cost a run to discover.
# api-level must be a POINT release. A bare 37 is not an SDK package and
# fails during setup, which cost a run to discover. `37.0` is the choice
# here rather than the only option: `37.1` and `37.2-beta*` exist and
# abort the same way, and api37-debug.yml's inputs document both, with
# the wrinkle that above 37.0 they ship only as google_apis_ps16k.
# docs/api-37-emulator-crash.md measures 37.0 rev 6 and 37.1 rev 8 side
# by side, so pinning 37.0 is a decision, not a constraint.
#
# notAnnotation removes the three tests that do not pass on this image; they
# run in the advisory job below, off the same marker so they cannot end up
@@ -355,6 +377,7 @@ jobs:
path: |
${{ runner.temp }}/logcat-api${{ matrix.label }}.txt
${{ runner.temp }}/diagnostics-api${{ matrix.label }}.txt
${{ runner.temp }}/gradle-api${{ matrix.label }}.txt
if-no-files-found: warn
# Only exists when the wrapper timeout tripped, so `ignore` keeps healthy runs quiet
@@ -445,6 +468,12 @@ jobs:
# The complement of the gating row's notAnnotation, off the same marker,
# so a test can never be excluded from both jobs or run in both.
E2E_EXTRA_GRADLE_ARGS: "-Pandroid.testInstrumentationRunnerArguments.annotation=org.libremediaconverter.FailsOnEmulatorApi37"
# Turns on the baseline comparison in the run-shape report, and only here. Every leg
# prints the shape; this is the one that also says whether it matches
# FAILS_ON_EMULATOR_API37_BASELINE, because this is the one whose test list is the
# marker. A deviation is a `::notice::` -- this job stays continue-on-error and stays
# out of the required contexts, so nothing the report finds can change a conclusion.
E2E_ADVISORY: "1"
with:
# Every device pin below matches the gating row exactly, so a difference
# between the two jobs is the test selection and nothing else.
@@ -477,6 +506,7 @@ jobs:
path: |
${{ runner.temp }}/logcat-api${{ env.E2E_LABEL }}.txt
${{ runner.temp }}/diagnostics-api${{ env.E2E_LABEL }}.txt
${{ runner.temp }}/gradle-api${{ env.E2E_LABEL }}.txt
if-no-files-found: warn
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
+15 -5
View File
@@ -89,6 +89,15 @@ days. Read it as the current answer, and see the git history if you need the old
not read a green run as evidence those three tests pass.
`docs/api-37-emulator-crash.md` has the measurements.
**That instruction is also why nobody looks, so the job now reports its own shape** — expected,
received, failed, and whether the run completed — to the job summary, and compares it against
`FAILS_ON_EMULATOR_API37_BASELINE`, committed beside the marker. A deviation is a `::notice::`;
the job stays advisory and its conclusion is untouched. **Add or remove a `@FailsOnEmulatorApi37`
and that number changes in the same diff**, or the next run says so. A bare failure count would
not have worked: the run is usually truncated by an `INSTRUMENTATION_ABORTED`, and the test XML
is written anyway and says nothing about it — `.github/scripts/e2e-report-shape.sh` is where that
is measured and explained.
Still true, and the reason the advisory job is not simply deleted: **API 37 needs a manual check on
the Pixel 10 Pro XL before each release.** Those three tests are the one thing CI cannot answer
for.
@@ -185,11 +194,12 @@ install for code that can never run — and on API 37 the full APK does not fit
`podman run --rm -v "$PWD:/mnt:z" docker.io/koalaman/shellcheck@sha256:61862eba... <files>`
(the digest is in `status_check.yml`; there is no shellcheck system package on this host).
**It does not cover inline `run:` blocks in the workflows**, and a good deal of this repo's bash
lives there. `actionlint` does cover them — it runs shellcheck over each `run:` — and reports one
pre-existing `info` finding in `build.yml`. It is not wired in because every action here is
pinned by SHA, and actionlint's usual installer is a `curl | bash` off a moving branch; doing it
properly means pinning a container digest. Tracked separately rather than bolted on.
**`actionlint` covers the half shellcheck cannot see** — the inline `run:` blocks, where a good
deal of this repo's bash lives. It runs shellcheck over each `run:` plus its own checks on
expression syntax, `needs:` references, matrix keys and action inputs. It sits in the same job,
**pinned by digest** for the reason above and one of its own: its documented installer is a
`curl | bash` off a moving branch, which does not belong in a repo that pins every action by SHA.
Locally: `podman run --rm -v "$PWD:/repo:z" -w /repo docker.io/rhysd/actionlint@sha256:9d360886... -color`.
## Dependency versions
+11
View File
@@ -1,3 +1,4 @@
import org.gradle.api.tasks.PathSensitivity
import org.gradle.testing.jacoco.tasks.JacocoReport
plugins {
@@ -206,6 +207,16 @@ detekt {
// `excludes` is not optional. Without it JaCoCo walks JDK-internal classes that Robolectric has
// no location for either, and the test JVM dies rather than reporting a number.
tasks.withType<Test>().configureEach {
// ReleasePermissionTest reads .github/workflows/build.yml, and Gradle cannot infer that a
// test depends on a file outside the source set. Without this the task stays UP-TO-DATE
// when the workflow changes, so the guard goes stale exactly when it matters. Measured:
// deleting the release job's `contents: write` and re-running gave "BUILD SUCCESSFUL in
// 614ms" with the test never executing; the same mutation under --rerun-tasks failed it.
// A guard that does not re-run when its subject changes is not a guard.
inputs.file(rootProject.file(".github/workflows/build.yml"))
.withPropertyName("releaseWorkflow")
.withPathSensitivity(PathSensitivity.RELATIVE)
extensions.configure<JacocoTaskExtension> {
isIncludeNoLocationClasses = true
excludes = listOf("jdk.internal.*")
@@ -18,7 +18,38 @@ package org.libremediaconverter
* Removing it is the goal, and the trigger is written down: a new API 37.x system image, or an
* ATD image for 37. Delete the annotation from the tests, and the advisory job goes empty and
* the gating one grows by two.
*
* **How many tests carry it is committed below**, as [FAILS_ON_EMULATOR_API37_BASELINE], and the
* advisory job checks the run against it. Adding or removing a marker means changing that number
* in the same diff.
*/
@Retention(AnnotationRetention.RUNTIME)
@Target(AnnotationTarget.CLASS, AnnotationTarget.FUNCTION)
annotation class FailsOnEmulatorApi37
/**
* How many tests carry [FailsOnEmulatorApi37] — the advisory API 37 job's committed baseline.
*
* **No Kotlin reads this, and it is not stray config.** `.github/scripts/e2e-report-shape.sh`
* parses it out of this file by name, with a line-anchored pattern, and the advisory job compares
* the run it just did against it: this many tests should start, and all of them should fail.
* Deleting it, renaming it, or indenting it into a class stops the comparison — the report would
* keep printing with nothing to compare to, so it announces that it could not read the baseline
* rather than falling quiet. If you see that notice, this line is what it means.
*
* **One number, both checks, and that is what the marker means.** A test carrying it cannot pass
* on this image, so the count is simultaneously how many the advisory leg runs and how many fail.
* A *smaller* failure count is the interesting direction: it means one of them now passes, which
* is the trigger the KDoc above names for deleting the annotation.
*
* So: adding or removing a [FailsOnEmulatorApi37] means changing this number, in this file, in
* the same diff. The report says so on the run itself if you forget — it prints the tree's own
* `grep` count beside this one.
*
* Why a baseline at all (#83): that job is `continue-on-error` and red on every PR by design, so
* a red X cannot distinguish the known failures from the known failures plus a new one. Counting
* failures alone does not fix it either — the run is usually truncated by an
* `INSTRUMENTATION_ABORTED`, so the count is a number taken from a partial run. The report
* records the truncation next to the counts for that reason.
*/
const val FAILS_ON_EMULATOR_API37_BASELINE = 3
@@ -36,8 +36,20 @@ import java.io.File
* 1. that the hardware path is worth having a second engine for at all, and
* 2. that x264's CRF is worth the GPL licence the app carries for it.
*
* Skips itself when the sample files are absent, so it is harmless in CI. Populate with:
* adb push <file>.mp4 /sdcard/Android/data/org.libremediaconverter/files/
* Skips itself when the sample files are absent, so it is harmless in CI — every green E2E
* leg reports two skips, and these are they.
*
* The two files it looks for, by exact name:
*
* - [H264_SAMPLE] for [hardwareVersusSoftwareOnRealVideo]
* - [AV1_SAMPLE] for [av1InputRoutesAccordingToDeviceDecodeSupport]
*
* **Where they go, and how, is on [samples] — read it before staging anything.** This used to
* carry an `adb push` line naming the external files dir, which [samples] then explains cannot
* work: a pushed file stays owned by the shell user and the app reads EACCES, surfacing as an
* unparseable input rather than a permission error. The instruction and its own refutation sat
* twelve lines apart. It is named in one place now rather than restated here, because restating
* it is what let the two drift.
*/
@UnstableApi
@RunWith(AndroidJUnit4::class)
@@ -0,0 +1,67 @@
package org.libremediaconverter.ci
import org.junit.Assert.assertTrue
import org.junit.Test
import java.io.File
/**
* That the release job still holds the one permission it needs to publish.
*
* `build.yml`'s `release` job declares `contents: write`, and nothing was checking it. Deleting
* those two lines leaves actionlint clean and CodeQL silent — a *narrower* permission is not an
* alert — and the job is `if: startsWith(github.ref, 'refs/tags/v')`, so no pull request and no
* merge to `main` can exercise it. Measured: with the declaration removed, every gating check
* still passes. The first thing that would notice is a release failing to publish, at the moment
* someone is trying to cut one.
*
* The deletion also looks like tidying. A top-level `permissions: contents: read` now sits
* directly above it, so a reader could reasonably take the job-level block for a duplicate. It is
* an override, not a duplicate, and a comment saying so is not a check.
*
* `BackupExclusionsTest` is the precedent: a file that is configuration rather than code, load
* bearing, and unguarded because nothing compiles it.
*
* **What this pins, and what it does not.** It asserts the declaration exists in the `release`
* job's block. It cannot assert that a release actually publishes — that needs a tag push, which
* is the thing no PR can do. So this is a tripwire against silent removal, not proof the release
* path works.
*/
class ReleasePermissionTest {
@Test
fun `the release job declares the write permission it needs to publish`() {
val release = jobBlock("release")
assertTrue(
"build.yml's `release` job no longer declares `contents: write`. It is the only " +
"permission that lets the job create a release, the top-level block above it is " +
"`contents: read`, and nothing else in CI would catch this until a tag failed to " +
"publish. If the release moved elsewhere, delete this test deliberately.",
release.any { it.trimStart().startsWith("contents: write") },
)
}
/**
* The lines of one top-level job, from its ` <name>:` header to the next job at that indent.
*
* Line-based rather than parsed: the module has no YAML dependency, and adding one to read two
* lines would be a worse trade than a scan that fails loudly when the shape changes.
*/
private fun jobBlock(name: String): List<String> {
val lines = workflow.readLines()
val start = lines.indexOfFirst { it == " $name:" }
check(start >= 0) { "no ` $name:` job in ${workflow.path} — has the file been restructured?" }
val rest = lines.drop(start + 1)
val end = rest.indexOfFirst { it.matches(Regex("^ {2}[A-Za-z0-9_-]+:.*")) }
return if (end < 0) rest else rest.take(end)
}
/**
* Found by walking up rather than by a fixed relative path: Gradle's working directory for the
* unit tests is the module, but that is a default rather than a promise.
*/
private val workflow: File
get() = generateSequence(File(".").absoluteFile) { it.parentFile }
.map { File(it, ".github/workflows/build.yml") }
.firstOrNull { it.isFile }
?: error("could not find .github/workflows/build.yml above ${File(".").absolutePath}")
}
@@ -140,10 +140,16 @@ class ConversionViewModelProbeFailureTest {
private fun pickedProbe(): InputProbe? {
val viewModel = ConversionViewModel(app, Dispatchers.Unconfined)
viewModel.onInputPicked(INPUT)
// The predicate is the guard, and it is the only one needed. It requires `Ready`, so a
// pick that ended in `Failed` never satisfies it and `awaitState` fails on its timeout
// naming what it was waiting for -- "Ready with a probe" -- which says more than a
// separate assertion could. A `ready as? ConversionState.Failed` check used to sit here
// and was dead: `Ready` and `Failed` are sibling subtypes of one sealed interface, so
// the cast was always null and the assertNull could never fire. Measured, not assumed --
// flipping it to assertNotNull failed all three callers of this helper.
val ready = awaitState(viewModel.state, "Ready with a probe") {
it is ConversionState.Ready && it.input.probe != null
}
assertNull("nothing here should reach a terminal failure", (ready as? ConversionState.Failed))
return (ready as ConversionState.Ready).input.probe
}
+9 -4
View File
@@ -32,8 +32,12 @@ reached* is not. Re-measured on 2026-08-22, seven runs, one variable at a time:
| r06 | `android-37.1` rev 8 | `swangle_indirect` | ANGLE | **yes, 285 s** | 23 |
| r07 | `android-37.0` rev 6 | `host` + `-feature -HostComposition` | host | **no**, wedged adb at 208 s | not readable |
The discriminator is exact across all seven: **a run boots if and only if the emulator log says
something other than `gles_mode_selected:host`.**
The discriminator is exact across the **six runs that reported**: a run boots if and only if the
emulator log says something other than `gles_mode_selected:host`. r07 is excluded on purpose — it
wedged adb at 208 s and is recorded below as inconclusive rather than ruled out, and a row this
page calls inconclusive cannot also be counted as evidence. Excluding it costs nothing: r07 is a
`host` row, so the discriminator predicts it would not boot, and confirming a prediction with the
one run whose evidence did not come back would add no information either way.
One caveat about how independent those rows are, because the table flatters itself. `-gpu
angle_indirect` (r05) and `-gpu swangle_indirect` (r03) both logged `gles_mode_selected:swangle`
@@ -509,8 +513,9 @@ API 37", not "is that codec broken".
`.github/workflows/api37-debug.yml` carried "roughly every 20 s" for the kill cycle in its own
comments. That number was the watchdog's **sampling** interval, not the cadence, and the two got
conflated. Measured across the seven runs above, gaps between successive `hasReadColorBufferDma`
aborts run **20 s to 90 s, median 60–70 s — three to five aborts in a four-minute window**.
conflated. Measured across the **six runs whose crash buffer could be read** — r07 wedged adb
before one could be taken, so it contributes no gaps — successive `hasReadColorBufferDma` aborts
run **20 s to 90 s, median 60–70 s — three to five aborts in a four-minute window**.
Slower than assumed, and still not slow enough: install, data-directory creation and
instrumentation start-up do not fit inside one gap.