Say what the advisory API 37 job actually found, so a new failure is not invisible #111

Merged
JMR-dev merged 4 commits from ci/advisory-failure-report into main 2026-08-26 02:00:58 +00:00
JMR-dev commented 2026-08-25 21:05:09 +00:00 (Migrated from github.com)

Closes #83.

The advisory API 37 job now ends by reporting what it actually found — expected, received,
failed, and whether the run completed — to $GITHUB_STEP_SUMMARY, and compares that against a
committed baseline.

The constraint is unchanged, deliberately

E2E API 37 Media3 hardware transcode (advisory) stays continue-on-error: true, stays red, and
stays out of ruleset 21117412's eight required contexts. A deviation is a ::notice::, never an
::error::. Nothing here can fail a build: e2e-report-shape.sh exits 0 unconditionally, every
field defaults to unknown, and every comparison is guarded — an unset variable under set -u is
exactly how a diagnostic becomes the thing that turns a leg red.

The ticket's "today" column is stale, and the correction matters

#83 was written from run 32800638011, where two of the three marked tests reported. I measured
nine advisory runs on 2026-08-25 rather than carrying that forward:

field #83 says measured today
expected 3 3
received 2 3 (test XML) / 2 (the truncation line)
failed 2 3
completed cleanly no usually no — 7 of 8 aborted, 1 did not

failed is 3, not 2, because AGP synthesises a failure for the test the abort truncated — all
three Execute …: FAILED lines are present in an aborted run. The abort itself is intermittent,
which is why completed cleanly is recorded but not compared: comparing it would announce a
deviation on a run that is fine. It is recorded because that is the field #108 would show up in on
a gating leg — and it does show up there, since every leg gets the shape.

The XML question, answered before building on it

#83 asked whether a test XML is even produced when the run aborts. It is — and that is worse
than absent: for run 32865281555, a run the runner had just described as truncated, the XML
reports a tidy <testsuites tests="3" failures="3"> and says nothing whatever about the abort. A
report that parsed only the XML would say "3 of 3, all accounted for" forever.

So the report reads both and prints which number came from where: the XML is the authority on how
many results landed, the runner's own output is the only authority on whether the run finished.
2>&1 | tee in e2e-run.sh, and the 2>&1 is load-bearing — the truncation line is not on
stdout, and capturing stdout alone would have left completed cleanly: yes permanently
unfalsifiable.

One number, beside the marker

FAILS_ON_EMULATOR_API37_BASELINE = 3, in FailsOnEmulatorApi37.kt. One number covers both
compared fields because that is what the marker means — a test carrying it cannot pass on this
image, so the count is simultaneously how many the advisory leg runs and how many fail. A smaller
failure count is the interesting direction: it means one now passes, which is the trigger the KDoc
already names for deleting the annotation. The report also prints #81's grep count beside the
baseline, so a stale baseline shows up on the run rather than only when the emulator disagrees.

The gating legs get the shape without the comparison — they run the whole suite, so comparing there
would announce a deviation five times a run.

Mutation, both directions

1. A fourth failing @FailsOnEmulatorApi37 test — run on real CI, from a throwaway branch
(#112, closed and deleted). Run 32899133552, advisory job 97968820644, verbatim:

Starting 4 tests on test(AVD) - 17
Tests on test(AVD) - 17 failed: There was 4 failure(s).
Test run failed to complete. Expected 4 tests, received 3. onError: commandError=false message=INSTRUMENTATION_ABORTED: System has crashed.
----- RUN SHAPE (api37-media3-transcode) -----
  expected:          4
  received:          4
  failed:            4
  completed cleanly: no
  received before the abort: 3
  failed tests:
    org.libremediaconverter.MutationProbeApi37Test.failsOnPurposeToProveTheAdvisoryReportFires
    org.libremediaconverter.convert.Media3EngineTest.runsFromAThreadWithNoLooper
    org.libremediaconverter.convert.Media3EngineTest.transcodesH264ToH265AndReportsProgress
    org.libremediaconverter.saf.SafPickerRoundTripTest.thePickedInputSurvivesARealRotation
  baseline DEVIATION: the runner started 4 tests, the baseline is 3
  baseline DEVIATION: 4 tests failed, the baseline is 3 — every test carrying the marker is expected to fail on this image, so fewer means one now passes and more means a new one joined
  baseline DEVIATION: the tree carries 4 tests marked `@FailsOnEmulatorApi37` but the baseline says 3 — update FAILS_ON_EMULATOR_API37_BASELINE
##[notice]E2E api37-media3-transcode: the runner started 4 tests, the baseline is 3
##[notice]E2E api37-media3-transcode: 4 tests failed, the baseline is 3 — ...
##[notice]E2E api37-media3-transcode: the tree carries 4 tests marked `@FailsOnEmulatorApi37` but the baseline says 3 — ...

Those three landed as real notice annotations on the check run, and the report names the new
test
, which is the thing a count alone cannot do. The advisory job's conclusion was failure
— the conclusion it has on every PR.
It stayed continue-on-error: true and stayed out of the
required contexts.

That run's run-level conclusion was also red, and that is the probe rather than the report:
the probe fails unconditionally, and @FailsOnEmulatorApi37 only removes it from the API 37
gating leg, so legs 33–36 ran it. API 33's own report says so — Starting 60 tests,
There was 1 failure(s), the one failure being MutationProbeApi37Test. The API 37 gating leg
stayed green.

2. The baseline number changed alone — run locally against the captured stdout of a real CI
run
, 32865281555, because the API 37 emulator image cannot be run on this host (CLAUDE.md:
local emulators are API 33–36). sed on the committed const, 3 → 4, nothing else touched:

----- RUN SHAPE (api37-media3-transcode) -----
  expected:          3
  received:          3
  failed:            3
  completed cleanly: no
  received before the abort: 2
  baseline DEVIATION: the runner started 3 tests, the baseline is 4
  baseline DEVIATION: 3 tests failed, the baseline is 4 — every test carrying the marker is expected to fail on this image, so fewer means one now passes and more means a new one joined
  baseline DEVIATION: the tree carries 3 tests marked `@FailsOnEmulatorApi37` but the baseline says 4 — update FAILS_ON_EMULATOR_API37_BASELINE

Restored to 3 immediately after; the same log against the committed baseline prints
baseline: matches (3 expected, 3 failed) and no notice.

3. An unreadable baseline is a deviation too, added after review pointed out that the
^-anchored sed means indenting the const into an object would silently switch the comparison
off while the table kept printing — the ticket's own failure mode wearing a different hat.
Verified locally three ways against the same real log: file absent → announces; const indented
into an object → announces; committed file unchanged → silent.

Degenerate shapes

A diagnostic that crashes, or that invents a number, is worse than none. Checked against real and
constructed logs: no log file, empty log file, Starting 0 tests (the framework-restart shape),
and a stale app/build XML with no test run in the log. All exit 0, none report a bogus 0 vs 3,
and the XML is read only when the runner said a run started — which matters on
tools/local-emulator/run-e2e.sh, where one checkout drives several API levels.

Not covered

The one thing I could not do on this host is run the API 37 leg locally, so mutation 2 and the
degenerate shapes were exercised against captured CI output rather than a live emulator.
Mutation 1 and the abort itself were seen on real CI, in run 32899133552
(Expected 4 tests, received 3 … INSTRUMENTATION_ABORTED), so the truncation capture is not
inferred.

The negative control, on this branch's own run

Run 32918773988, advisory job 98028007674 — the same code, nothing mutated:

----- RUN SHAPE (api37-media3-transcode) -----
  expected:          3
  received:          3
  failed:            3
  completed cleanly: no
  received before the abort: 2
  failed tests:
    org.libremediaconverter.convert.Media3EngineTest.runsFromAThreadWithNoLooper
    org.libremediaconverter.convert.Media3EngineTest.transcodesH264ToH265AndReportsProgress
    org.libremediaconverter.saf.SafPickerRoundTripTest.thePickedInputSurvivesARealRotation
  baseline: matches (3 expected, 3 failed)
  (the table above is also on the job summary page)

Zero notice annotations on that job; the mutation run had three. That pair is the
demonstration — silent when it matches, loud when it does not. The last line is there because
GitHub's check-run API returns summary: null for a job summary, so a write that silently did not
happen would be invisible; it confirms GITHUB_STEP_SUMMARY reaches the emulator-runner's child
process.

E2E API 33 was red on that run with one failure,
ReattachOnLaunchTest.doesNotOverwriteAPickTheUserHasAlreadyMade — that is #49, a known flake, and
the report's own failed tests: list is how it was identified, without opening a log. Re-running
the failed legs put the run at conclusion success: E2E API 33 green, all eight required
contexts green, and the advisory job on that attempt (98029549102) reporting
expected 3 / received 3 / failed 3 / completed cleanly: no, baseline: matches, zero notices.

The run's conclusion is success while the advisory job's conclusion is failure. That is the
constraint holding, demonstrated rather than asserted: the report ran, compared, found a match, and
the advisory red still did not gate anything.

Closes #83. The advisory API 37 job now ends by reporting **what it actually found** — expected, received, failed, and whether the run completed — to `$GITHUB_STEP_SUMMARY`, and compares that against a committed baseline. ### The constraint is unchanged, deliberately `E2E API 37 Media3 hardware transcode (advisory)` stays `continue-on-error: true`, stays red, and stays out of ruleset `21117412`'s eight required contexts. A deviation is a `::notice::`, never an `::error::`. Nothing here can fail a build: `e2e-report-shape.sh` exits 0 unconditionally, every field defaults to `unknown`, and every comparison is guarded — an unset variable under `set -u` is exactly how a diagnostic becomes the thing that turns a leg red. ### The ticket's "today" column is stale, and the correction matters #83 was written from run `32800638011`, where two of the three marked tests reported. I measured nine advisory runs on 2026-08-25 rather than carrying that forward: | field | #83 says | measured today | |---|---|---| | expected | 3 | 3 | | received | 2 | 3 (test XML) / 2 (the truncation line) | | failed | 2 | **3** | | completed cleanly | no | **usually no** — 7 of 8 aborted, 1 did not | `failed` is 3, not 2, because AGP synthesises a failure for the test the abort truncated — all three `Execute …: FAILED` lines are present in an aborted run. The abort itself is *intermittent*, which is why `completed cleanly` is **recorded but not compared**: comparing it would announce a deviation on a run that is fine. It is recorded because that is the field #108 would show up in on a gating leg — and it does show up there, since every leg gets the shape. ### The XML question, answered before building on it #83 asked whether a test XML is even produced when the run aborts. **It is** — and that is worse than absent: for run `32865281555`, a run the runner had just described as truncated, the XML reports a tidy `<testsuites tests="3" failures="3">` and says nothing whatever about the abort. A report that parsed only the XML would say "3 of 3, all accounted for" forever. So the report reads both and prints which number came from where: the XML is the authority on how many results landed, the runner's own output is the only authority on whether the run finished. `2>&1 | tee` in `e2e-run.sh`, and the `2>&1` is load-bearing — the truncation line is not on stdout, and capturing stdout alone would have left `completed cleanly: yes` permanently unfalsifiable. ### One number, beside the marker `FAILS_ON_EMULATOR_API37_BASELINE = 3`, in `FailsOnEmulatorApi37.kt`. One number covers both compared fields because that is what the marker means — a test carrying it cannot pass on this image, so the count is simultaneously how many the advisory leg runs and how many fail. A *smaller* failure count is the interesting direction: it means one now passes, which is the trigger the KDoc already names for deleting the annotation. The report also prints #81's `grep` count beside the baseline, so a stale baseline shows up on the run rather than only when the emulator disagrees. The gating legs get the shape without the comparison — they run the whole suite, so comparing there would announce a deviation five times a run. ### Mutation, both directions **1. A fourth failing `@FailsOnEmulatorApi37` test — run on real CI**, from a throwaway branch (#112, closed and deleted). Run `32899133552`, advisory job `97968820644`, verbatim: ``` Starting 4 tests on test(AVD) - 17 Tests on test(AVD) - 17 failed: There was 4 failure(s). Test run failed to complete. Expected 4 tests, received 3. onError: commandError=false message=INSTRUMENTATION_ABORTED: System has crashed. ----- RUN SHAPE (api37-media3-transcode) ----- expected: 4 received: 4 failed: 4 completed cleanly: no received before the abort: 3 failed tests: org.libremediaconverter.MutationProbeApi37Test.failsOnPurposeToProveTheAdvisoryReportFires org.libremediaconverter.convert.Media3EngineTest.runsFromAThreadWithNoLooper org.libremediaconverter.convert.Media3EngineTest.transcodesH264ToH265AndReportsProgress org.libremediaconverter.saf.SafPickerRoundTripTest.thePickedInputSurvivesARealRotation baseline DEVIATION: the runner started 4 tests, the baseline is 3 baseline DEVIATION: 4 tests failed, the baseline is 3 — every test carrying the marker is expected to fail on this image, so fewer means one now passes and more means a new one joined baseline DEVIATION: the tree carries 4 tests marked `@FailsOnEmulatorApi37` but the baseline says 3 — update FAILS_ON_EMULATOR_API37_BASELINE ##[notice]E2E api37-media3-transcode: the runner started 4 tests, the baseline is 3 ##[notice]E2E api37-media3-transcode: 4 tests failed, the baseline is 3 — ... ##[notice]E2E api37-media3-transcode: the tree carries 4 tests marked `@FailsOnEmulatorApi37` but the baseline says 3 — ... ``` Those three landed as real `notice` annotations on the check run, and the report **names the new test**, which is the thing a count alone cannot do. **The advisory job's conclusion was `failure` — the conclusion it has on every PR.** It stayed `continue-on-error: true` and stayed out of the required contexts. That run's *run-level* conclusion was also red, and that is the probe rather than the report: the probe fails unconditionally, and `@FailsOnEmulatorApi37` only removes it from the **API 37** gating leg, so legs 33–36 ran it. API 33's own report says so — `Starting 60 tests`, `There was 1 failure(s)`, the one failure being `MutationProbeApi37Test`. The API 37 gating leg stayed green. **2. The baseline number changed alone — run locally against the captured stdout of a real CI run**, `32865281555`, because the API 37 emulator image cannot be run on this host (CLAUDE.md: local emulators are API 33–36). `sed` on the committed const, `3` → `4`, nothing else touched: ``` ----- RUN SHAPE (api37-media3-transcode) ----- expected: 3 received: 3 failed: 3 completed cleanly: no received before the abort: 2 baseline DEVIATION: the runner started 3 tests, the baseline is 4 baseline DEVIATION: 3 tests failed, the baseline is 4 — every test carrying the marker is expected to fail on this image, so fewer means one now passes and more means a new one joined baseline DEVIATION: the tree carries 3 tests marked `@FailsOnEmulatorApi37` but the baseline says 4 — update FAILS_ON_EMULATOR_API37_BASELINE ``` Restored to `3` immediately after; the same log against the committed baseline prints `baseline: matches (3 expected, 3 failed)` and no notice. **3. An unreadable baseline is a deviation too**, added after review pointed out that the `^`-anchored `sed` means indenting the const into an object would silently switch the comparison off while the table kept printing — the ticket's own failure mode wearing a different hat. Verified locally three ways against the same real log: file absent → announces; const indented into an `object` → announces; committed file unchanged → silent. ### Degenerate shapes A diagnostic that crashes, or that invents a number, is worse than none. Checked against real and constructed logs: no log file, empty log file, `Starting 0 tests` (the framework-restart shape), and a stale `app/build` XML with no test run in the log. All exit 0, none report a bogus `0 vs 3`, and the XML is read only when the runner said a run started — which matters on `tools/local-emulator/run-e2e.sh`, where one checkout drives several API levels. ### Not covered The one thing I could not do on this host is run the API 37 leg locally, so mutation 2 and the degenerate shapes were exercised against **captured** CI output rather than a live emulator. Mutation 1 and the abort itself were seen on real CI, in run `32899133552` (`Expected 4 tests, received 3` … `INSTRUMENTATION_ABORTED`), so the truncation capture is not inferred. ### The negative control, on this branch's own run Run `32918773988`, advisory job `98028007674` — the same code, nothing mutated: ``` ----- RUN SHAPE (api37-media3-transcode) ----- expected: 3 received: 3 failed: 3 completed cleanly: no received before the abort: 2 failed tests: org.libremediaconverter.convert.Media3EngineTest.runsFromAThreadWithNoLooper org.libremediaconverter.convert.Media3EngineTest.transcodesH264ToH265AndReportsProgress org.libremediaconverter.saf.SafPickerRoundTripTest.thePickedInputSurvivesARealRotation baseline: matches (3 expected, 3 failed) (the table above is also on the job summary page) ``` **Zero `notice` annotations on that job; the mutation run had three.** That pair is the demonstration — silent when it matches, loud when it does not. The last line is there because GitHub's check-run API returns `summary: null` for a job summary, so a write that silently did not happen would be invisible; it confirms `GITHUB_STEP_SUMMARY` reaches the emulator-runner's child process. `E2E API 33` was red on that run with one failure, `ReattachOnLaunchTest.doesNotOverwriteAPickTheUserHasAlreadyMade` — that is #49, a known flake, and the report's own `failed tests:` list is how it was identified, without opening a log. Re-running the failed legs put the run at **conclusion `success`**: `E2E API 33` green, all eight required contexts green, and the advisory job on that attempt (`98029549102`) reporting `expected 3 / received 3 / failed 3 / completed cleanly: no`, `baseline: matches`, zero notices. **The run's conclusion is `success` while the advisory job's conclusion is `failure`.** That is the constraint holding, demonstrated rather than asserted: the report ran, compared, found a match, and the advisory red still did not gate anything.
Sign in to join this conversation.