Compare commits
20
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
bf2214a549 | ||
|
|
989069207e | ||
|
|
72ff7adfcc | ||
|
|
994ea8a3dd | ||
|
|
3e9528454c | ||
|
|
0702916229 | ||
|
|
3fd34c24a6 | ||
|
|
d01a46a708 | ||
|
|
3806641cb2 | ||
|
|
5f9498150c | ||
|
|
8d8703ab49 | ||
|
|
27b7654418 | ||
|
|
4a8e30099e | ||
|
|
e7d84cc69f | ||
|
|
a8b494b846 | ||
|
|
865a4a7c8e | ||
|
|
e856679395 | ||
|
|
49c483d877 | ||
|
|
1b220856ab | ||
|
|
d37c391c60 |
Executable
+294
@@ -0,0 +1,294 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# Reports the SHAPE of an instrumented run -- how many tests were expected, how many
|
||||
# reported, how many failed, and whether the run completed at all -- to the step log and to
|
||||
# the job summary. In advisory mode it also compares that shape against a committed baseline
|
||||
# and says plainly whether it matches.
|
||||
#
|
||||
# WHY THIS EXISTS (#83): the advisory API 37 leg is red on every PR by design, so a NEW failure
|
||||
# joining the known ones is invisible -- nothing in a red X distinguishes "the known ones" from
|
||||
# "the known ones plus yours". CLAUDE.md tells everyone not to read that job's red as their
|
||||
# change breaking something, which is correct, and which also means nobody looks.
|
||||
#
|
||||
# WHY NOT A BARE FAILURE COUNT, measured rather than assumed. On this image the run is usually
|
||||
# truncated: `Test run failed to complete. Expected 3 tests, received 2.` with
|
||||
# `INSTRUMENTATION_ABORTED: System has crashed.` A count taken from a truncated run misleads in
|
||||
# both directions -- a fourth marked test can still yield the same number if the abort lands
|
||||
# earlier, and the known set getting worse can LOWER it. So all four fields are recorded, and
|
||||
# the one saying the run was truncated is recorded with them.
|
||||
#
|
||||
# WHY IT IS A SEPARATE SCRIPT rather than a function inside e2e-run.sh: it is a pure seam. It
|
||||
# reads a captured log plus the test XML and writes a report, so it can be run against a REAL
|
||||
# log saved from a REAL CI run -- which is how the baseline comparison was shown to fire
|
||||
# without waiting on an emulator. `git ls-files '*.sh'` also picks it up for shellcheck for
|
||||
# free.
|
||||
#
|
||||
# THIS SCRIPT NEVER FAILS A RUN. It is a diagnostic, and e2e-run.sh's header explains why that
|
||||
# rule is absolute here. Every field defaults to `unknown` and every comparison is guarded,
|
||||
# because an unset variable under `set -u`, or a `[ "" -eq 3 ]`, is exactly how a diagnostic
|
||||
# becomes the thing that turns a leg red. It exits 0 unconditionally.
|
||||
#
|
||||
# Usage:
|
||||
# e2e-report-shape.sh <label> <gradle-log> [<baseline-file>]
|
||||
#
|
||||
# With a third argument the run is compared against the baseline in that file (advisory mode)
|
||||
# and a `::notice::` is emitted per deviation. NEVER `::error::`: the advisory job is
|
||||
# `continue-on-error: true` and stays that way, and an error annotation would be a new way for
|
||||
# a diagnostic to change a conclusion.
|
||||
set -uo pipefail
|
||||
|
||||
LABEL="${1:-unknown}"
|
||||
LOG="${2:-}"
|
||||
BASELINE_FILE="${3:-}"
|
||||
|
||||
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
|
||||
REPO_ROOT="$(cd -- "$SCRIPT_DIR/../.." && pwd)"
|
||||
XML_DIR="$REPO_ROOT/app/build/outputs/androidTest-results/connected/debug"
|
||||
|
||||
# Gradle colours its output even when it is piped, so `FAILED` arrives wrapped in escape codes.
|
||||
# The numeric lines parsed below are not coloured, but stripping is cheap insurance against a
|
||||
# pattern that would otherwise silently match nothing.
|
||||
ESC="$(printf '\033')"
|
||||
scan() { [ -s "$LOG" ] && sed -e "s/${ESC}\[[0-9;]*[a-zA-Z]//g" -- "$LOG"; }
|
||||
first_number() { grep -oE '[0-9]+' | head -1; }
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Source 1: the runner's own output. This is the ONLY place a truncation is visible. The test
|
||||
# XML read below is written even for an aborted run and says nothing whatever about the abort
|
||||
# -- measured on run 32865281555, where the XML reports a tidy tests="3" failures="3" for a run
|
||||
# the runner had just described as truncated. That is the reason this parses stdout at all.
|
||||
# ---------------------------------------------------------------------------
|
||||
starting_line="$(scan | grep -aoE 'Starting [0-9]+ tests on .*' | tail -1)"
|
||||
abort_line="$(scan | grep -aoE 'Test run failed to complete\. Expected [0-9]+ tests, received [0-9]+\.' | tail -1)"
|
||||
aborted_hits="$(scan | grep -ac 'INSTRUMENTATION_ABORTED' || true)"
|
||||
failure_line="$(scan | grep -aoE 'There was [0-9]+ failure\(s\)\.' | tail -1)"
|
||||
failed_names="$(scan | grep -aoE 'Execute [A-Za-z0-9_.$]+: FAILED' | sed -e 's/^Execute //' -e 's/: FAILED$//' | sort -u)"
|
||||
|
||||
expected="$(printf '%s' "$starting_line" | first_number)"
|
||||
expected_src="\`$starting_line\`"
|
||||
abort_expected="$(printf '%s' "$abort_line" | grep -oE 'Expected [0-9]+' | first_number)"
|
||||
abort_received="$(printf '%s' "$abort_line" | grep -oE 'received [0-9]+' | first_number)"
|
||||
log_failed="$(printf '%s' "$failure_line" | first_number)"
|
||||
|
||||
# `Starting N tests` is missing when the framework restarted under the run and Gradle never got
|
||||
# a test list. The truncation line still carries the number it was told to expect.
|
||||
if [ -z "$expected" ] && [ -n "$abort_expected" ]; then
|
||||
expected="$abort_expected"
|
||||
expected_src="\`$abort_line\`"
|
||||
fi
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Source 2: the JUnit XML. Measured on both a truncated advisory run and a green gating leg:
|
||||
# `<testsuites tests="N" failures="M">` is present in both, and aggregates every suite. It is
|
||||
# the authority on how many results landed and how many were failures. It is NOT an authority
|
||||
# on whether the run finished, which is what source 1 is for.
|
||||
# ---------------------------------------------------------------------------
|
||||
#
|
||||
# Read ONLY when the runner said a test run happened. `app/build` survives between runs on a
|
||||
# developer machine -- tools/local-emulator/run-e2e.sh drives several API levels against one
|
||||
# checkout -- so a leg that never got as far as starting tests would otherwise be reported from
|
||||
# the previous leg's XML, which is a wrong answer rather than a missing one.
|
||||
xml_head=""
|
||||
xml_count=0
|
||||
if [ -n "$starting_line$abort_line" ] && [ -d "$XML_DIR" ]; then
|
||||
while IFS= read -r f; do
|
||||
xml_count=$((xml_count + 1))
|
||||
[ -z "$xml_head" ] && xml_head="$(grep -ao '<testsuites[^>]*>' "$f" | head -1)"
|
||||
done < <(find "$XML_DIR" -maxdepth 1 -name 'TEST-*.xml' -print 2> /dev/null | sort)
|
||||
fi
|
||||
xml_tests="$(printf '%s' "$xml_head" | grep -oE ' tests="[0-9]+"' | first_number)"
|
||||
xml_failed="$(printf '%s' "$xml_head" | grep -oE ' failures="[0-9]+"' | first_number)"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Derive the four fields, each with where its number came from. Everything stays a string, so a
|
||||
# missing source reads `unknown` rather than becoming 0 -- a report claiming 0 tests when it
|
||||
# merely could not see them would announce a deviation on every cancelled run.
|
||||
# ---------------------------------------------------------------------------
|
||||
received="unknown"
|
||||
received_src="no source"
|
||||
if [ -n "$xml_tests" ]; then
|
||||
received="$xml_tests"
|
||||
received_src="test XML \`<testsuites tests=\"$xml_tests\">\`"
|
||||
elif [ -n "$abort_received" ]; then
|
||||
received="$abort_received"
|
||||
received_src="\`$abort_line\`"
|
||||
elif [ -n "$expected" ] && [ -z "$abort_line" ]; then
|
||||
received="$expected"
|
||||
received_src="the run was not truncated, so every expected test reported"
|
||||
fi
|
||||
|
||||
failed="unknown"
|
||||
failed_src="no source"
|
||||
if [ -n "$xml_failed" ]; then
|
||||
failed="$xml_failed"
|
||||
failed_src="test XML \`<testsuites failures=\"$xml_failed\">\`"
|
||||
elif [ -n "$log_failed" ]; then
|
||||
failed="$log_failed"
|
||||
failed_src="\`$failure_line\`"
|
||||
fi
|
||||
|
||||
if [ -z "$expected" ]; then
|
||||
expected="unknown"
|
||||
expected_src="no \`Starting N tests\` line"
|
||||
fi
|
||||
|
||||
# A run whose start nobody can see is not a run of zero tests. Cancellation (this workflow sets
|
||||
# cancel-in-progress) and the `Starting 0 tests` shape a framework restart produces both land
|
||||
# here, and both have to say so rather than compare a number that does not exist.
|
||||
no_run="none"
|
||||
if [ "$expected" = "unknown" ] && [ "$received" = "unknown" ]; then
|
||||
no_run="nothing"
|
||||
elif [ "$expected" = "0" ]; then
|
||||
no_run="zero"
|
||||
fi
|
||||
|
||||
if [ -n "$abort_line" ]; then
|
||||
completed="**no**"
|
||||
completed_src="\`$abort_line\` with \`INSTRUMENTATION_ABORTED\`"
|
||||
elif [ "${aborted_hits:-0}" -gt 0 ]; then
|
||||
completed="**no**"
|
||||
completed_src="\`INSTRUMENTATION_ABORTED\` in the runner output"
|
||||
elif [ "$no_run" = "nothing" ]; then
|
||||
# "cleanly" would be a lie about a run that left no evidence it happened.
|
||||
completed="unknown"
|
||||
completed_src="no runner output to read"
|
||||
else
|
||||
completed="yes"
|
||||
completed_src="no truncation line and no \`INSTRUMENTATION_ABORTED\`"
|
||||
fi
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Advisory mode: compare against the committed baseline.
|
||||
#
|
||||
# ONE number covers both compared fields, and that is deliberate rather than a shortcut. The
|
||||
# marker means "cannot pass on this image", so the number of tests carrying it is both how many
|
||||
# the advisory leg should run and how many should fail. A smaller `failed` means one now passes
|
||||
# -- which is the trigger to delete the annotation, written down in FailsOnEmulatorApi37.kt.
|
||||
# ---------------------------------------------------------------------------
|
||||
#
|
||||
# `advisory` and `baseline` are two variables on purpose. "A comparison was asked for" and "a
|
||||
# number was found to compare against" are different facts, and collapsing them is how this
|
||||
# report would go quietly back to being the thing #83 filed: the `sed` below is anchored, so
|
||||
# indenting the const into an object -- or renaming it, or moving it to another file -- empties
|
||||
# `baseline`, and a single flag would take the whole comparison down with it while the table
|
||||
# kept printing. An unreadable baseline is itself a deviation, and is announced as one.
|
||||
advisory="no"
|
||||
baseline=""
|
||||
marked=""
|
||||
deviations=()
|
||||
if [ -n "$BASELINE_FILE" ]; then
|
||||
advisory="yes"
|
||||
[ -f "$BASELINE_FILE" ] \
|
||||
&& baseline="$(sed -nE 's/^const val FAILS_ON_EMULATOR_API37_BASELINE = ([0-9]+).*/\1/p' "$BASELINE_FILE" | head -1)"
|
||||
# The #81 check, verbatim: what the tree actually carries. Reported next to the baseline so a
|
||||
# stale baseline shows up here rather than only once the emulator disagrees with it.
|
||||
if [ -d "$REPO_ROOT/app/src/androidTest" ]; then
|
||||
marked="$(grep -rn "@FailsOnEmulatorApi37" "$REPO_ROOT/app/src/androidTest" --include='*.kt' \
|
||||
| grep -v import | grep -c FailsOn || true)"
|
||||
fi
|
||||
fi
|
||||
|
||||
if [ "$advisory" = "yes" ] && [ -z "$baseline" ]; then
|
||||
deviations+=("the committed baseline could not be read from \`$(basename -- "$BASELINE_FILE")\` — has \`FAILS_ON_EMULATOR_API37_BASELINE\` been renamed, indented into a class, or moved? Nothing was compared")
|
||||
fi
|
||||
|
||||
if [ -n "$baseline" ]; then
|
||||
if [ "$no_run" = "nothing" ]; then
|
||||
deviations+=("no test run observed — the runner never reported starting one, where the baseline expects $baseline tests carrying \`@FailsOnEmulatorApi37\`")
|
||||
elif [ "$no_run" = "zero" ]; then
|
||||
deviations+=("the runner started 0 tests, where the baseline expects $baseline — on this image that is the framework having restarted under the run, not an empty test list")
|
||||
else
|
||||
if [ "$expected" != "unknown" ] && [ "$expected" != "$baseline" ]; then
|
||||
deviations+=("the runner started $expected tests, the baseline is $baseline")
|
||||
fi
|
||||
if [ "$failed" != "unknown" ] && [ "$failed" != "$baseline" ]; then
|
||||
deviations+=("$failed tests failed, the baseline is $baseline — every test carrying the marker is expected to fail on this image, so fewer means one now passes and more means a new one joined")
|
||||
fi
|
||||
fi
|
||||
if [ -n "$marked" ] && [ "$marked" != "$baseline" ]; then
|
||||
deviations+=("the tree carries $marked tests marked \`@FailsOnEmulatorApi37\` but the baseline says $baseline — update FAILS_ON_EMULATOR_API37_BASELINE")
|
||||
fi
|
||||
fi
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Emit. Step log first, so the common case needs neither the summary page nor an artifact.
|
||||
# ---------------------------------------------------------------------------
|
||||
echo "----- RUN SHAPE (api${LABEL}) -----"
|
||||
echo " expected: $expected"
|
||||
echo " received: $received"
|
||||
echo " failed: $failed"
|
||||
echo " completed cleanly: ${completed//\*/}"
|
||||
if [ -n "$abort_received" ]; then
|
||||
echo " received before the abort: $abort_received"
|
||||
fi
|
||||
if [ -n "$failed_names" ]; then
|
||||
echo " failed tests:"
|
||||
printf '%s\n' "$failed_names" | sed -e 's/^/ /'
|
||||
fi
|
||||
if [ "$advisory" = "yes" ]; then
|
||||
if [ "${#deviations[@]}" -eq 0 ]; then
|
||||
echo " baseline: matches ($baseline expected, $baseline failed)"
|
||||
else
|
||||
printf ' baseline DEVIATION: %s\n' "${deviations[@]}"
|
||||
fi
|
||||
fi
|
||||
|
||||
# A notice, never an error. See the header.
|
||||
if [ "${#deviations[@]}" -gt 0 ]; then
|
||||
for d in "${deviations[@]}"; do
|
||||
echo "::notice::E2E api${LABEL}: $d"
|
||||
done
|
||||
fi
|
||||
|
||||
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
{
|
||||
echo "### E2E api${LABEL} — run shape"
|
||||
echo
|
||||
echo "| field | value | where it came from |"
|
||||
echo "| --- | --- | --- |"
|
||||
echo "| expected | $expected | $expected_src |"
|
||||
echo "| received | $received | $received_src |"
|
||||
echo "| failed | $failed | $failed_src |"
|
||||
echo "| completed cleanly | $completed | $completed_src |"
|
||||
if [ -n "$abort_received" ]; then
|
||||
echo "| received before the abort | $abort_received | the same line — the XML above counts the truncated test as a failure, this number does not |"
|
||||
fi
|
||||
echo
|
||||
if [ -n "$failed_names" ]; then
|
||||
echo "Failed:"
|
||||
echo
|
||||
printf '%s\n' "$failed_names" | sed -e 's/^/- `/' -e 's/$/`/'
|
||||
echo
|
||||
fi
|
||||
if [ "$xml_count" -gt 1 ]; then
|
||||
echo "> $xml_count test XML files were present; the counts above come from the first."
|
||||
echo
|
||||
fi
|
||||
if [ "$advisory" = "yes" ]; then
|
||||
if [ "${#deviations[@]}" -eq 0 ]; then
|
||||
echo "**Matches the committed baseline of $baseline** — $baseline tests carry \`@FailsOnEmulatorApi37\` and all $baseline failed, which is what this job is for."
|
||||
elif [ -z "$baseline" ]; then
|
||||
echo "**The committed baseline could not be read, so nothing was compared.** Announced as a notice, not an error: this job is advisory and its conclusion is unchanged by anything here."
|
||||
echo
|
||||
printf -- '- %s\n' "${deviations[@]}"
|
||||
else
|
||||
echo "**DEVIATION from the committed baseline of $baseline.** Announced as a notice, not an error: this job is advisory and its conclusion is unchanged by anything here."
|
||||
echo
|
||||
printf -- '- %s\n' "${deviations[@]}"
|
||||
fi
|
||||
echo
|
||||
echo "<sub>The baseline lives beside the marker, in \`FailsOnEmulatorApi37.kt\`. \`completed cleanly\` is recorded rather than compared: the truncation is intermittent — of eight advisory runs read on 2026-08-25, seven aborted and one did not — so comparing it would announce a deviation on a run that is fine.</sub>"
|
||||
else
|
||||
echo "<sub>No baseline comparison: that is the advisory API 37 leg only. The shape is recorded here anyway because a truncated run reports fewer results than it ran, which is what issue #108 looks like on a gating leg.</sub>"
|
||||
fi
|
||||
echo
|
||||
} >> "$GITHUB_STEP_SUMMARY"
|
||||
# The summary page is the deliverable -- "readable without opening a log" is what #83 asked
|
||||
# for -- and GitHub exposes no API for reading a job summary back, so a write that silently
|
||||
# did not happen would be invisible. This line is in the step log, which can be read.
|
||||
echo " (the table above is also on the job summary page)"
|
||||
else
|
||||
echo " (GITHUB_STEP_SUMMARY is unset -- step log only)"
|
||||
fi
|
||||
|
||||
exit 0
|
||||
@@ -34,6 +34,13 @@ TMP="${RUNNER_TEMP:-/tmp}"
|
||||
LOGCAT_LOG="$TMP/logcat-api${LABEL}.txt"
|
||||
DIAG_LOG="$TMP/diagnostics-api${LABEL}.txt"
|
||||
WEDGE_LOG="$TMP/wedge-diagnostics-api${LABEL}.txt"
|
||||
# Gradle's own output, captured to a file as well as the step log, because the run-shape report
|
||||
# below has to parse it. Uploaded with the diagnostics, so a report that reads wrong can be
|
||||
# checked against what it read.
|
||||
GRADLE_LOG="$TMP/gradle-api${LABEL}.txt"
|
||||
|
||||
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
|
||||
REPO_ROOT="$(cd -- "$SCRIPT_DIR/../.." && pwd)"
|
||||
|
||||
# ~5 min is a healthy leg (measured across API 33-36), and this wraps only the gradle client,
|
||||
# a subset of that. 20 min is generous enough never to trip on a slow-but-working run, and far
|
||||
@@ -214,12 +221,40 @@ status=0
|
||||
# script rather than forking it: that runs several API levels back to back against one checkout
|
||||
# and passes `--rerun`, so a level cannot be skipped as up-to-date and report the previous
|
||||
# level's results as its own. CI gets a fresh runner per level and does not need it.
|
||||
#
|
||||
# `2>&1 | tee`, and the `2>&1` is the load-bearing half. The step log merges both streams, so
|
||||
# reading one cannot tell you which stream a line came from -- and the single line the report
|
||||
# below needs most, `Test run failed to complete. ... INSTRUMENTATION_ABORTED`, is not on
|
||||
# stdout. Capturing stdout alone would leave the report saying "completed cleanly: yes" forever,
|
||||
# which is precisely the comparison that cannot fire. pipefail is already set and tee exits 0,
|
||||
# so the pipeline's status is still gradle's -- including the 124 that means the wrapper fired.
|
||||
#
|
||||
# `tee` and not `tee -a`, unlike the logcat above: CI gets a fresh runner per leg, but
|
||||
# tools/local-emulator/run-e2e.sh reuses one machine, and an appended log would have the report
|
||||
# reading the PREVIOUS run of the same API level. The console goes plain rather than showing
|
||||
# gradle's live progress bar, which is what it already did in CI.
|
||||
# shellcheck disable=SC2086
|
||||
timeout -k 30s "$WEDGE_TIMEOUT" \
|
||||
./gradlew :app:connectedDebugAndroidTest -PabiFilters=x86_64 --stacktrace \
|
||||
${E2E_EXTRA_GRADLE_ARGS:-} || status=$?
|
||||
${E2E_EXTRA_GRADLE_ARGS:-} 2>&1 | tee "$GRADLE_LOG" || status=$?
|
||||
echo "::endgroup::"
|
||||
|
||||
# The run-shape report: expected/received/failed and whether the run finished, every time,
|
||||
# green or red. It never changes `status` -- it is a diagnostic, and the header's rule about
|
||||
# diagnostics applies to it as much as to every probe below.
|
||||
#
|
||||
# The baseline argument, and only it, turns on the comparison, and only the advisory API 37 job
|
||||
# passes E2E_ADVISORY=1. Comparing on the gating legs would announce a deviation on all five of
|
||||
# them every run, since they run the whole suite rather than the marked three. They still get
|
||||
# the report: a truncated run reporting fewer results than it ran is what #108 looks like, and
|
||||
# `completed cleanly` is the field that shows it.
|
||||
if [ "${E2E_ADVISORY:-}" = "1" ]; then
|
||||
bash "$SCRIPT_DIR/e2e-report-shape.sh" "$LABEL" "$GRADLE_LOG" \
|
||||
"$REPO_ROOT/app/src/androidTest/java/org/libremediaconverter/FailsOnEmulatorApi37.kt" || true
|
||||
else
|
||||
bash "$SCRIPT_DIR/e2e-report-shape.sh" "$LABEL" "$GRADLE_LOG" || true
|
||||
fi
|
||||
|
||||
if [ "$status" -eq 0 ]; then
|
||||
kill "$LOGCAT_PID" 2>/dev/null || true
|
||||
exit 0
|
||||
|
||||
@@ -13,6 +13,17 @@ on:
|
||||
# reference amounts to running whatever that repository contains tomorrow. This matters
|
||||
# more here than on pull requests: these jobs sign nothing today, but they do publish
|
||||
# the artifacts people install.
|
||||
# Declared here rather than inherited, for the reason status_check.yml gives for its own
|
||||
# block: the token's reach should be readable in the file that uses it, and a repository
|
||||
# default that widens later should not silently widen these jobs with it. The repository
|
||||
# default is `read` today, so this changes nothing about what runs -- it fixes what a
|
||||
# reader can know without leaving the file, and it is what CodeQL alert #1 asked for.
|
||||
#
|
||||
# The `release` job below overrides this with `contents: write`, which is how job-level
|
||||
# permissions work: this is a default, not a ceiling.
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
env:
|
||||
GRADLE_CACHE_PATHS: |
|
||||
~/.gradle/caches
|
||||
|
||||
@@ -280,8 +280,13 @@ jobs:
|
||||
# docs/api-37-emulator-crash.md has the per-method measurements, and the
|
||||
# correction that produced them.
|
||||
#
|
||||
# api-level must be "37.0". A bare 37 is not an SDK package and fails
|
||||
# during setup, which cost a run to discover.
|
||||
# api-level must be a POINT release. A bare 37 is not an SDK package and
|
||||
# fails during setup, which cost a run to discover. `37.0` is the choice
|
||||
# here rather than the only option: `37.1` and `37.2-beta*` exist and
|
||||
# abort the same way, and api37-debug.yml's inputs document both, with
|
||||
# the wrinkle that above 37.0 they ship only as google_apis_ps16k.
|
||||
# docs/api-37-emulator-crash.md measures 37.0 rev 6 and 37.1 rev 8 side
|
||||
# by side, so pinning 37.0 is a decision, not a constraint.
|
||||
#
|
||||
# notAnnotation removes the three tests that do not pass on this image; they
|
||||
# run in the advisory job below, off the same marker so they cannot end up
|
||||
@@ -372,6 +377,7 @@ jobs:
|
||||
path: |
|
||||
${{ runner.temp }}/logcat-api${{ matrix.label }}.txt
|
||||
${{ runner.temp }}/diagnostics-api${{ matrix.label }}.txt
|
||||
${{ runner.temp }}/gradle-api${{ matrix.label }}.txt
|
||||
if-no-files-found: warn
|
||||
|
||||
# Only exists when the wrapper timeout tripped, so `ignore` keeps healthy runs quiet
|
||||
@@ -462,6 +468,12 @@ jobs:
|
||||
# The complement of the gating row's notAnnotation, off the same marker,
|
||||
# so a test can never be excluded from both jobs or run in both.
|
||||
E2E_EXTRA_GRADLE_ARGS: "-Pandroid.testInstrumentationRunnerArguments.annotation=org.libremediaconverter.FailsOnEmulatorApi37"
|
||||
# Turns on the baseline comparison in the run-shape report, and only here. Every leg
|
||||
# prints the shape; this is the one that also says whether it matches
|
||||
# FAILS_ON_EMULATOR_API37_BASELINE, because this is the one whose test list is the
|
||||
# marker. A deviation is a `::notice::` -- this job stays continue-on-error and stays
|
||||
# out of the required contexts, so nothing the report finds can change a conclusion.
|
||||
E2E_ADVISORY: "1"
|
||||
with:
|
||||
# Every device pin below matches the gating row exactly, so a difference
|
||||
# between the two jobs is the test selection and nothing else.
|
||||
@@ -494,6 +506,7 @@ jobs:
|
||||
path: |
|
||||
${{ runner.temp }}/logcat-api${{ env.E2E_LABEL }}.txt
|
||||
${{ runner.temp }}/diagnostics-api${{ env.E2E_LABEL }}.txt
|
||||
${{ runner.temp }}/gradle-api${{ env.E2E_LABEL }}.txt
|
||||
if-no-files-found: warn
|
||||
|
||||
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
||||
|
||||
@@ -89,6 +89,15 @@ days. Read it as the current answer, and see the git history if you need the old
|
||||
not read a green run as evidence those three tests pass.
|
||||
`docs/api-37-emulator-crash.md` has the measurements.
|
||||
|
||||
**That instruction is also why nobody looks, so the job now reports its own shape** — expected,
|
||||
received, failed, and whether the run completed — to the job summary, and compares it against
|
||||
`FAILS_ON_EMULATOR_API37_BASELINE`, committed beside the marker. A deviation is a `::notice::`;
|
||||
the job stays advisory and its conclusion is untouched. **Add or remove a `@FailsOnEmulatorApi37`
|
||||
and that number changes in the same diff**, or the next run says so. A bare failure count would
|
||||
not have worked: the run is usually truncated by an `INSTRUMENTATION_ABORTED`, and the test XML
|
||||
is written anyway and says nothing about it — `.github/scripts/e2e-report-shape.sh` is where that
|
||||
is measured and explained.
|
||||
|
||||
Still true, and the reason the advisory job is not simply deleted: **API 37 needs a manual check on
|
||||
the Pixel 10 Pro XL before each release.** Those three tests are the one thing CI cannot answer
|
||||
for.
|
||||
|
||||
@@ -1,3 +1,4 @@
|
||||
import org.gradle.api.tasks.PathSensitivity
|
||||
import org.gradle.testing.jacoco.tasks.JacocoReport
|
||||
|
||||
plugins {
|
||||
@@ -206,6 +207,16 @@ detekt {
|
||||
// `excludes` is not optional. Without it JaCoCo walks JDK-internal classes that Robolectric has
|
||||
// no location for either, and the test JVM dies rather than reporting a number.
|
||||
tasks.withType<Test>().configureEach {
|
||||
// ReleasePermissionTest reads .github/workflows/build.yml, and Gradle cannot infer that a
|
||||
// test depends on a file outside the source set. Without this the task stays UP-TO-DATE
|
||||
// when the workflow changes, so the guard goes stale exactly when it matters. Measured:
|
||||
// deleting the release job's `contents: write` and re-running gave "BUILD SUCCESSFUL in
|
||||
// 614ms" with the test never executing; the same mutation under --rerun-tasks failed it.
|
||||
// A guard that does not re-run when its subject changes is not a guard.
|
||||
inputs.file(rootProject.file(".github/workflows/build.yml"))
|
||||
.withPropertyName("releaseWorkflow")
|
||||
.withPathSensitivity(PathSensitivity.RELATIVE)
|
||||
|
||||
extensions.configure<JacocoTaskExtension> {
|
||||
isIncludeNoLocationClasses = true
|
||||
excludes = listOf("jdk.internal.*")
|
||||
|
||||
@@ -18,7 +18,38 @@ package org.libremediaconverter
|
||||
* Removing it is the goal, and the trigger is written down: a new API 37.x system image, or an
|
||||
* ATD image for 37. Delete the annotation from the tests, and the advisory job goes empty and
|
||||
* the gating one grows by two.
|
||||
*
|
||||
* **How many tests carry it is committed below**, as [FAILS_ON_EMULATOR_API37_BASELINE], and the
|
||||
* advisory job checks the run against it. Adding or removing a marker means changing that number
|
||||
* in the same diff.
|
||||
*/
|
||||
@Retention(AnnotationRetention.RUNTIME)
|
||||
@Target(AnnotationTarget.CLASS, AnnotationTarget.FUNCTION)
|
||||
annotation class FailsOnEmulatorApi37
|
||||
|
||||
/**
|
||||
* How many tests carry [FailsOnEmulatorApi37] — the advisory API 37 job's committed baseline.
|
||||
*
|
||||
* **No Kotlin reads this, and it is not stray config.** `.github/scripts/e2e-report-shape.sh`
|
||||
* parses it out of this file by name, with a line-anchored pattern, and the advisory job compares
|
||||
* the run it just did against it: this many tests should start, and all of them should fail.
|
||||
* Deleting it, renaming it, or indenting it into a class stops the comparison — the report would
|
||||
* keep printing with nothing to compare to, so it announces that it could not read the baseline
|
||||
* rather than falling quiet. If you see that notice, this line is what it means.
|
||||
*
|
||||
* **One number, both checks, and that is what the marker means.** A test carrying it cannot pass
|
||||
* on this image, so the count is simultaneously how many the advisory leg runs and how many fail.
|
||||
* A *smaller* failure count is the interesting direction: it means one of them now passes, which
|
||||
* is the trigger the KDoc above names for deleting the annotation.
|
||||
*
|
||||
* So: adding or removing a [FailsOnEmulatorApi37] means changing this number, in this file, in
|
||||
* the same diff. The report says so on the run itself if you forget — it prints the tree's own
|
||||
* `grep` count beside this one.
|
||||
*
|
||||
* Why a baseline at all (#83): that job is `continue-on-error` and red on every PR by design, so
|
||||
* a red X cannot distinguish the known failures from the known failures plus a new one. Counting
|
||||
* failures alone does not fix it either — the run is usually truncated by an
|
||||
* `INSTRUMENTATION_ABORTED`, so the count is a number taken from a partial run. The report
|
||||
* records the truncation next to the counts for that reason.
|
||||
*/
|
||||
const val FAILS_ON_EMULATOR_API37_BASELINE = 3
|
||||
|
||||
@@ -36,8 +36,20 @@ import java.io.File
|
||||
* 1. that the hardware path is worth having a second engine for at all, and
|
||||
* 2. that x264's CRF is worth the GPL licence the app carries for it.
|
||||
*
|
||||
* Skips itself when the sample files are absent, so it is harmless in CI. Populate with:
|
||||
* adb push <file>.mp4 /sdcard/Android/data/org.libremediaconverter/files/
|
||||
* Skips itself when the sample files are absent, so it is harmless in CI — every green E2E
|
||||
* leg reports two skips, and these are they.
|
||||
*
|
||||
* The two files it looks for, by exact name:
|
||||
*
|
||||
* - [H264_SAMPLE] for [hardwareVersusSoftwareOnRealVideo]
|
||||
* - [AV1_SAMPLE] for [av1InputRoutesAccordingToDeviceDecodeSupport]
|
||||
*
|
||||
* **Where they go, and how, is on [samples] — read it before staging anything.** This used to
|
||||
* carry an `adb push` line naming the external files dir, which [samples] then explains cannot
|
||||
* work: a pushed file stays owned by the shell user and the app reads EACCES, surfacing as an
|
||||
* unparseable input rather than a permission error. The instruction and its own refutation sat
|
||||
* twelve lines apart. It is named in one place now rather than restated here, because restating
|
||||
* it is what let the two drift.
|
||||
*/
|
||||
@UnstableApi
|
||||
@RunWith(AndroidJUnit4::class)
|
||||
|
||||
@@ -0,0 +1,67 @@
|
||||
package org.libremediaconverter.ci
|
||||
|
||||
import org.junit.Assert.assertTrue
|
||||
import org.junit.Test
|
||||
import java.io.File
|
||||
|
||||
/**
|
||||
* That the release job still holds the one permission it needs to publish.
|
||||
*
|
||||
* `build.yml`'s `release` job declares `contents: write`, and nothing was checking it. Deleting
|
||||
* those two lines leaves actionlint clean and CodeQL silent — a *narrower* permission is not an
|
||||
* alert — and the job is `if: startsWith(github.ref, 'refs/tags/v')`, so no pull request and no
|
||||
* merge to `main` can exercise it. Measured: with the declaration removed, every gating check
|
||||
* still passes. The first thing that would notice is a release failing to publish, at the moment
|
||||
* someone is trying to cut one.
|
||||
*
|
||||
* The deletion also looks like tidying. A top-level `permissions: contents: read` now sits
|
||||
* directly above it, so a reader could reasonably take the job-level block for a duplicate. It is
|
||||
* an override, not a duplicate, and a comment saying so is not a check.
|
||||
*
|
||||
* `BackupExclusionsTest` is the precedent: a file that is configuration rather than code, load
|
||||
* bearing, and unguarded because nothing compiles it.
|
||||
*
|
||||
* **What this pins, and what it does not.** It asserts the declaration exists in the `release`
|
||||
* job's block. It cannot assert that a release actually publishes — that needs a tag push, which
|
||||
* is the thing no PR can do. So this is a tripwire against silent removal, not proof the release
|
||||
* path works.
|
||||
*/
|
||||
class ReleasePermissionTest {
|
||||
|
||||
@Test
|
||||
fun `the release job declares the write permission it needs to publish`() {
|
||||
val release = jobBlock("release")
|
||||
assertTrue(
|
||||
"build.yml's `release` job no longer declares `contents: write`. It is the only " +
|
||||
"permission that lets the job create a release, the top-level block above it is " +
|
||||
"`contents: read`, and nothing else in CI would catch this until a tag failed to " +
|
||||
"publish. If the release moved elsewhere, delete this test deliberately.",
|
||||
release.any { it.trimStart().startsWith("contents: write") },
|
||||
)
|
||||
}
|
||||
|
||||
/**
|
||||
* The lines of one top-level job, from its ` <name>:` header to the next job at that indent.
|
||||
*
|
||||
* Line-based rather than parsed: the module has no YAML dependency, and adding one to read two
|
||||
* lines would be a worse trade than a scan that fails loudly when the shape changes.
|
||||
*/
|
||||
private fun jobBlock(name: String): List<String> {
|
||||
val lines = workflow.readLines()
|
||||
val start = lines.indexOfFirst { it == " $name:" }
|
||||
check(start >= 0) { "no ` $name:` job in ${workflow.path} — has the file been restructured?" }
|
||||
val rest = lines.drop(start + 1)
|
||||
val end = rest.indexOfFirst { it.matches(Regex("^ {2}[A-Za-z0-9_-]+:.*")) }
|
||||
return if (end < 0) rest else rest.take(end)
|
||||
}
|
||||
|
||||
/**
|
||||
* Found by walking up rather than by a fixed relative path: Gradle's working directory for the
|
||||
* unit tests is the module, but that is a default rather than a promise.
|
||||
*/
|
||||
private val workflow: File
|
||||
get() = generateSequence(File(".").absoluteFile) { it.parentFile }
|
||||
.map { File(it, ".github/workflows/build.yml") }
|
||||
.firstOrNull { it.isFile }
|
||||
?: error("could not find .github/workflows/build.yml above ${File(".").absolutePath}")
|
||||
}
|
||||
@@ -32,8 +32,12 @@ reached* is not. Re-measured on 2026-08-22, seven runs, one variable at a time:
|
||||
| r06 | `android-37.1` rev 8 | `swangle_indirect` | ANGLE | **yes, 285 s** | 23 |
|
||||
| r07 | `android-37.0` rev 6 | `host` + `-feature -HostComposition` | host | **no**, wedged adb at 208 s | not readable |
|
||||
|
||||
The discriminator is exact across all seven: **a run boots if and only if the emulator log says
|
||||
something other than `gles_mode_selected:host`.**
|
||||
The discriminator is exact across the **six runs that reported**: a run boots if and only if the
|
||||
emulator log says something other than `gles_mode_selected:host`. r07 is excluded on purpose — it
|
||||
wedged adb at 208 s and is recorded below as inconclusive rather than ruled out, and a row this
|
||||
page calls inconclusive cannot also be counted as evidence. Excluding it costs nothing: r07 is a
|
||||
`host` row, so the discriminator predicts it would not boot, and confirming a prediction with the
|
||||
one run whose evidence did not come back would add no information either way.
|
||||
|
||||
One caveat about how independent those rows are, because the table flatters itself. `-gpu
|
||||
angle_indirect` (r05) and `-gpu swangle_indirect` (r03) both logged `gles_mode_selected:swangle`
|
||||
@@ -509,8 +513,9 @@ API 37", not "is that codec broken".
|
||||
|
||||
`.github/workflows/api37-debug.yml` carried "roughly every 20 s" for the kill cycle in its own
|
||||
comments. That number was the watchdog's **sampling** interval, not the cadence, and the two got
|
||||
conflated. Measured across the seven runs above, gaps between successive `hasReadColorBufferDma`
|
||||
aborts run **20 s to 90 s, median 60–70 s — three to five aborts in a four-minute window**.
|
||||
conflated. Measured across the **six runs whose crash buffer could be read** — r07 wedged adb
|
||||
before one could be taken, so it contributes no gaps — successive `hasReadColorBufferDma` aborts
|
||||
run **20 s to 90 s, median 60–70 s — three to five aborts in a four-minute window**.
|
||||
Slower than assumed, and still not slow enough: install, data-directory creation and
|
||||
instrumentation start-up do not fit inside one gap.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user