The advisory API 37 job now ends by reporting what it actually found — expected, received,
failed, and whether the run completed — to $GITHUB_STEP_SUMMARY, and compares that against a
committed baseline.
The constraint is unchanged, deliberately
E2E API 37 Media3 hardware transcode (advisory) stays continue-on-error: true, stays red, and
stays out of ruleset 21117412's eight required contexts. A deviation is a ::notice::, never an ::error::. Nothing here can fail a build: e2e-report-shape.sh exits 0 unconditionally, every
field defaults to unknown, and every comparison is guarded — an unset variable under set -u is
exactly how a diagnostic becomes the thing that turns a leg red.
The ticket's "today" column is stale, and the correction matters
#83 was written from run 32800638011, where two of the three marked tests reported. I measured
nine advisory runs on 2026-08-25 rather than carrying that forward:
failed is 3, not 2, because AGP synthesises a failure for the test the abort truncated — all
three Execute …: FAILED lines are present in an aborted run. The abort itself is intermittent,
which is why completed cleanly is recorded but not compared: comparing it would announce a
deviation on a run that is fine. It is recorded because that is the field #108 would show up in on
a gating leg — and it does show up there, since every leg gets the shape.
The XML question, answered before building on it
#83 asked whether a test XML is even produced when the run aborts. It is — and that is worse
than absent: for run 32865281555, a run the runner had just described as truncated, the XML
reports a tidy <testsuites tests="3" failures="3"> and says nothing whatever about the abort. A
report that parsed only the XML would say "3 of 3, all accounted for" forever.
So the report reads both and prints which number came from where: the XML is the authority on how
many results landed, the runner's own output is the only authority on whether the run finished. 2>&1 | tee in e2e-run.sh, and the 2>&1 is load-bearing — the truncation line is not on
stdout, and capturing stdout alone would have left completed cleanly: yes permanently
unfalsifiable.
One number, beside the marker
FAILS_ON_EMULATOR_API37_BASELINE = 3, in FailsOnEmulatorApi37.kt. One number covers both
compared fields because that is what the marker means — a test carrying it cannot pass on this
image, so the count is simultaneously how many the advisory leg runs and how many fail. A smaller
failure count is the interesting direction: it means one now passes, which is the trigger the KDoc
already names for deleting the annotation. The report also prints #81's grep count beside the
baseline, so a stale baseline shows up on the run rather than only when the emulator disagrees.
The gating legs get the shape without the comparison — they run the whole suite, so comparing there
would announce a deviation five times a run.
Mutation, both directions
1. A fourth failing @FailsOnEmulatorApi37 test — run on real CI, from a throwaway branch
(#112, closed and deleted). Run 32899133552, advisory job 97968820644, verbatim:
Starting 4 tests on test(AVD) - 17
Tests on test(AVD) - 17 failed: There was 4 failure(s).
Test run failed to complete. Expected 4 tests, received 3. onError: commandError=false message=INSTRUMENTATION_ABORTED: System has crashed.
----- RUN SHAPE (api37-media3-transcode) -----
expected: 4
received: 4
failed: 4
completed cleanly: no
received before the abort: 3
failed tests:
org.libremediaconverter.MutationProbeApi37Test.failsOnPurposeToProveTheAdvisoryReportFires
org.libremediaconverter.convert.Media3EngineTest.runsFromAThreadWithNoLooper
org.libremediaconverter.convert.Media3EngineTest.transcodesH264ToH265AndReportsProgress
org.libremediaconverter.saf.SafPickerRoundTripTest.thePickedInputSurvivesARealRotation
baseline DEVIATION: the runner started 4 tests, the baseline is 3
baseline DEVIATION: 4 tests failed, the baseline is 3 — every test carrying the marker is expected to fail on this image, so fewer means one now passes and more means a new one joined
baseline DEVIATION: the tree carries 4 tests marked `@FailsOnEmulatorApi37` but the baseline says 3 — update FAILS_ON_EMULATOR_API37_BASELINE
##[notice]E2E api37-media3-transcode: the runner started 4 tests, the baseline is 3
##[notice]E2E api37-media3-transcode: 4 tests failed, the baseline is 3 — ...
##[notice]E2E api37-media3-transcode: the tree carries 4 tests marked `@FailsOnEmulatorApi37` but the baseline says 3 — ...
Those three landed as real notice annotations on the check run, and the report names the new
test, which is the thing a count alone cannot do. The advisory job's conclusion was failure
— the conclusion it has on every PR. It stayed continue-on-error: true and stayed out of the
required contexts.
That run's run-level conclusion was also red, and that is the probe rather than the report:
the probe fails unconditionally, and @FailsOnEmulatorApi37 only removes it from the API 37
gating leg, so legs 33–36 ran it. API 33's own report says so — Starting 60 tests, There was 1 failure(s), the one failure being MutationProbeApi37Test. The API 37 gating leg
stayed green.
2. The baseline number changed alone — run locally against the captured stdout of a real CI
run, 32865281555, because the API 37 emulator image cannot be run on this host (CLAUDE.md:
local emulators are API 33–36). sed on the committed const, 3 → 4, nothing else touched:
----- RUN SHAPE (api37-media3-transcode) -----
expected: 3
received: 3
failed: 3
completed cleanly: no
received before the abort: 2
baseline DEVIATION: the runner started 3 tests, the baseline is 4
baseline DEVIATION: 3 tests failed, the baseline is 4 — every test carrying the marker is expected to fail on this image, so fewer means one now passes and more means a new one joined
baseline DEVIATION: the tree carries 3 tests marked `@FailsOnEmulatorApi37` but the baseline says 4 — update FAILS_ON_EMULATOR_API37_BASELINE
Restored to 3 immediately after; the same log against the committed baseline prints baseline: matches (3 expected, 3 failed) and no notice.
3. An unreadable baseline is a deviation too, added after review pointed out that the ^-anchored sed means indenting the const into an object would silently switch the comparison
off while the table kept printing — the ticket's own failure mode wearing a different hat.
Verified locally three ways against the same real log: file absent → announces; const indented
into an object → announces; committed file unchanged → silent.
Degenerate shapes
A diagnostic that crashes, or that invents a number, is worse than none. Checked against real and
constructed logs: no log file, empty log file, Starting 0 tests (the framework-restart shape),
and a stale app/build XML with no test run in the log. All exit 0, none report a bogus 0 vs 3,
and the XML is read only when the runner said a run started — which matters on tools/local-emulator/run-e2e.sh, where one checkout drives several API levels.
Not covered
The one thing I could not do on this host is run the API 37 leg locally, so mutation 2 and the
degenerate shapes were exercised against captured CI output rather than a live emulator.
Mutation 1 and the abort itself were seen on real CI, in run 32899133552
(Expected 4 tests, received 3 … INSTRUMENTATION_ABORTED), so the truncation capture is not
inferred.
The negative control, on this branch's own run
Run 32918773988, advisory job 98028007674 — the same code, nothing mutated:
----- RUN SHAPE (api37-media3-transcode) -----
expected: 3
received: 3
failed: 3
completed cleanly: no
received before the abort: 2
failed tests:
org.libremediaconverter.convert.Media3EngineTest.runsFromAThreadWithNoLooper
org.libremediaconverter.convert.Media3EngineTest.transcodesH264ToH265AndReportsProgress
org.libremediaconverter.saf.SafPickerRoundTripTest.thePickedInputSurvivesARealRotation
baseline: matches (3 expected, 3 failed)
(the table above is also on the job summary page)
Zero notice annotations on that job; the mutation run had three. That pair is the
demonstration — silent when it matches, loud when it does not. The last line is there because
GitHub's check-run API returns summary: null for a job summary, so a write that silently did not
happen would be invisible; it confirms GITHUB_STEP_SUMMARY reaches the emulator-runner's child
process.
E2E API 33 was red on that run with one failure, ReattachOnLaunchTest.doesNotOverwriteAPickTheUserHasAlreadyMade — that is #49, a known flake, and
the report's own failed tests: list is how it was identified, without opening a log. Re-running
the failed legs put the run at conclusion success: E2E API 33 green, all eight required
contexts green, and the advisory job on that attempt (98029549102) reporting expected 3 / received 3 / failed 3 / completed cleanly: no, baseline: matches, zero notices.
The run's conclusion is success while the advisory job's conclusion is failure. That is the
constraint holding, demonstrated rather than asserted: the report ran, compared, found a match, and
the advisory red still did not gate anything.
Closes #83.
The advisory API 37 job now ends by reporting **what it actually found** — expected, received,
failed, and whether the run completed — to `$GITHUB_STEP_SUMMARY`, and compares that against a
committed baseline.
### The constraint is unchanged, deliberately
`E2E API 37 Media3 hardware transcode (advisory)` stays `continue-on-error: true`, stays red, and
stays out of ruleset `21117412`'s eight required contexts. A deviation is a `::notice::`, never an
`::error::`. Nothing here can fail a build: `e2e-report-shape.sh` exits 0 unconditionally, every
field defaults to `unknown`, and every comparison is guarded — an unset variable under `set -u` is
exactly how a diagnostic becomes the thing that turns a leg red.
### The ticket's "today" column is stale, and the correction matters
#83 was written from run `32800638011`, where two of the three marked tests reported. I measured
nine advisory runs on 2026-08-25 rather than carrying that forward:
| field | #83 says | measured today |
|---|---|---|
| expected | 3 | 3 |
| received | 2 | 3 (test XML) / 2 (the truncation line) |
| failed | 2 | **3** |
| completed cleanly | no | **usually no** — 7 of 8 aborted, 1 did not |
`failed` is 3, not 2, because AGP synthesises a failure for the test the abort truncated — all
three `Execute …: FAILED` lines are present in an aborted run. The abort itself is *intermittent*,
which is why `completed cleanly` is **recorded but not compared**: comparing it would announce a
deviation on a run that is fine. It is recorded because that is the field #108 would show up in on
a gating leg — and it does show up there, since every leg gets the shape.
### The XML question, answered before building on it
#83 asked whether a test XML is even produced when the run aborts. **It is** — and that is worse
than absent: for run `32865281555`, a run the runner had just described as truncated, the XML
reports a tidy `<testsuites tests="3" failures="3">` and says nothing whatever about the abort. A
report that parsed only the XML would say "3 of 3, all accounted for" forever.
So the report reads both and prints which number came from where: the XML is the authority on how
many results landed, the runner's own output is the only authority on whether the run finished.
`2>&1 | tee` in `e2e-run.sh`, and the `2>&1` is load-bearing — the truncation line is not on
stdout, and capturing stdout alone would have left `completed cleanly: yes` permanently
unfalsifiable.
### One number, beside the marker
`FAILS_ON_EMULATOR_API37_BASELINE = 3`, in `FailsOnEmulatorApi37.kt`. One number covers both
compared fields because that is what the marker means — a test carrying it cannot pass on this
image, so the count is simultaneously how many the advisory leg runs and how many fail. A *smaller*
failure count is the interesting direction: it means one now passes, which is the trigger the KDoc
already names for deleting the annotation. The report also prints #81's `grep` count beside the
baseline, so a stale baseline shows up on the run rather than only when the emulator disagrees.
The gating legs get the shape without the comparison — they run the whole suite, so comparing there
would announce a deviation five times a run.
### Mutation, both directions
**1. A fourth failing `@FailsOnEmulatorApi37` test — run on real CI**, from a throwaway branch
(#112, closed and deleted). Run `32899133552`, advisory job `97968820644`, verbatim:
```
Starting 4 tests on test(AVD) - 17
Tests on test(AVD) - 17 failed: There was 4 failure(s).
Test run failed to complete. Expected 4 tests, received 3. onError: commandError=false message=INSTRUMENTATION_ABORTED: System has crashed.
----- RUN SHAPE (api37-media3-transcode) -----
expected: 4
received: 4
failed: 4
completed cleanly: no
received before the abort: 3
failed tests:
org.libremediaconverter.MutationProbeApi37Test.failsOnPurposeToProveTheAdvisoryReportFires
org.libremediaconverter.convert.Media3EngineTest.runsFromAThreadWithNoLooper
org.libremediaconverter.convert.Media3EngineTest.transcodesH264ToH265AndReportsProgress
org.libremediaconverter.saf.SafPickerRoundTripTest.thePickedInputSurvivesARealRotation
baseline DEVIATION: the runner started 4 tests, the baseline is 3
baseline DEVIATION: 4 tests failed, the baseline is 3 — every test carrying the marker is expected to fail on this image, so fewer means one now passes and more means a new one joined
baseline DEVIATION: the tree carries 4 tests marked `@FailsOnEmulatorApi37` but the baseline says 3 — update FAILS_ON_EMULATOR_API37_BASELINE
##[notice]E2E api37-media3-transcode: the runner started 4 tests, the baseline is 3
##[notice]E2E api37-media3-transcode: 4 tests failed, the baseline is 3 — ...
##[notice]E2E api37-media3-transcode: the tree carries 4 tests marked `@FailsOnEmulatorApi37` but the baseline says 3 — ...
```
Those three landed as real `notice` annotations on the check run, and the report **names the new
test**, which is the thing a count alone cannot do. **The advisory job's conclusion was `failure`
— the conclusion it has on every PR.** It stayed `continue-on-error: true` and stayed out of the
required contexts.
That run's *run-level* conclusion was also red, and that is the probe rather than the report:
the probe fails unconditionally, and `@FailsOnEmulatorApi37` only removes it from the **API 37**
gating leg, so legs 33–36 ran it. API 33's own report says so — `Starting 60 tests`,
`There was 1 failure(s)`, the one failure being `MutationProbeApi37Test`. The API 37 gating leg
stayed green.
**2. The baseline number changed alone — run locally against the captured stdout of a real CI
run**, `32865281555`, because the API 37 emulator image cannot be run on this host (CLAUDE.md:
local emulators are API 33–36). `sed` on the committed const, `3` → `4`, nothing else touched:
```
----- RUN SHAPE (api37-media3-transcode) -----
expected: 3
received: 3
failed: 3
completed cleanly: no
received before the abort: 2
baseline DEVIATION: the runner started 3 tests, the baseline is 4
baseline DEVIATION: 3 tests failed, the baseline is 4 — every test carrying the marker is expected to fail on this image, so fewer means one now passes and more means a new one joined
baseline DEVIATION: the tree carries 3 tests marked `@FailsOnEmulatorApi37` but the baseline says 4 — update FAILS_ON_EMULATOR_API37_BASELINE
```
Restored to `3` immediately after; the same log against the committed baseline prints
`baseline: matches (3 expected, 3 failed)` and no notice.
**3. An unreadable baseline is a deviation too**, added after review pointed out that the
`^`-anchored `sed` means indenting the const into an object would silently switch the comparison
off while the table kept printing — the ticket's own failure mode wearing a different hat.
Verified locally three ways against the same real log: file absent → announces; const indented
into an `object` → announces; committed file unchanged → silent.
### Degenerate shapes
A diagnostic that crashes, or that invents a number, is worse than none. Checked against real and
constructed logs: no log file, empty log file, `Starting 0 tests` (the framework-restart shape),
and a stale `app/build` XML with no test run in the log. All exit 0, none report a bogus `0 vs 3`,
and the XML is read only when the runner said a run started — which matters on
`tools/local-emulator/run-e2e.sh`, where one checkout drives several API levels.
### Not covered
The one thing I could not do on this host is run the API 37 leg locally, so mutation 2 and the
degenerate shapes were exercised against **captured** CI output rather than a live emulator.
Mutation 1 and the abort itself were seen on real CI, in run `32899133552`
(`Expected 4 tests, received 3` … `INSTRUMENTATION_ABORTED`), so the truncation capture is not
inferred.
### The negative control, on this branch's own run
Run `32918773988`, advisory job `98028007674` — the same code, nothing mutated:
```
----- RUN SHAPE (api37-media3-transcode) -----
expected: 3
received: 3
failed: 3
completed cleanly: no
received before the abort: 2
failed tests:
org.libremediaconverter.convert.Media3EngineTest.runsFromAThreadWithNoLooper
org.libremediaconverter.convert.Media3EngineTest.transcodesH264ToH265AndReportsProgress
org.libremediaconverter.saf.SafPickerRoundTripTest.thePickedInputSurvivesARealRotation
baseline: matches (3 expected, 3 failed)
(the table above is also on the job summary page)
```
**Zero `notice` annotations on that job; the mutation run had three.** That pair is the
demonstration — silent when it matches, loud when it does not. The last line is there because
GitHub's check-run API returns `summary: null` for a job summary, so a write that silently did not
happen would be invisible; it confirms `GITHUB_STEP_SUMMARY` reaches the emulator-runner's child
process.
`E2E API 33` was red on that run with one failure,
`ReattachOnLaunchTest.doesNotOverwriteAPickTheUserHasAlreadyMade` — that is #49, a known flake, and
the report's own `failed tests:` list is how it was identified, without opening a log. Re-running
the failed legs put the run at **conclusion `success`**: `E2E API 33` green, all eight required
contexts green, and the advisory job on that attempt (`98029549102`) reporting
`expected 3 / received 3 / failed 3 / completed cleanly: no`, `baseline: matches`, zero notices.
**The run's conclusion is `success` while the advisory job's conclusion is `failure`.** That is the
constraint holding, demonstrated rather than asserted: the report ran, compared, found a match, and
the advisory red still did not gate anything.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Closes #83.
The advisory API 37 job now ends by reporting what it actually found — expected, received,
failed, and whether the run completed — to
$GITHUB_STEP_SUMMARY, and compares that against acommitted baseline.
The constraint is unchanged, deliberately
E2E API 37 Media3 hardware transcode (advisory)stayscontinue-on-error: true, stays red, andstays out of ruleset
21117412's eight required contexts. A deviation is a::notice::, never an::error::. Nothing here can fail a build:e2e-report-shape.shexits 0 unconditionally, everyfield defaults to
unknown, and every comparison is guarded — an unset variable underset -uisexactly how a diagnostic becomes the thing that turns a leg red.
The ticket's "today" column is stale, and the correction matters
#83 was written from run
32800638011, where two of the three marked tests reported. I measurednine advisory runs on 2026-08-25 rather than carrying that forward:
failedis 3, not 2, because AGP synthesises a failure for the test the abort truncated — allthree
Execute …: FAILEDlines are present in an aborted run. The abort itself is intermittent,which is why
completed cleanlyis recorded but not compared: comparing it would announce adeviation on a run that is fine. It is recorded because that is the field #108 would show up in on
a gating leg — and it does show up there, since every leg gets the shape.
The XML question, answered before building on it
#83 asked whether a test XML is even produced when the run aborts. It is — and that is worse
than absent: for run
32865281555, a run the runner had just described as truncated, the XMLreports a tidy
<testsuites tests="3" failures="3">and says nothing whatever about the abort. Areport that parsed only the XML would say "3 of 3, all accounted for" forever.
So the report reads both and prints which number came from where: the XML is the authority on how
many results landed, the runner's own output is the only authority on whether the run finished.
2>&1 | teeine2e-run.sh, and the2>&1is load-bearing — the truncation line is not onstdout, and capturing stdout alone would have left
completed cleanly: yespermanentlyunfalsifiable.
One number, beside the marker
FAILS_ON_EMULATOR_API37_BASELINE = 3, inFailsOnEmulatorApi37.kt. One number covers bothcompared fields because that is what the marker means — a test carrying it cannot pass on this
image, so the count is simultaneously how many the advisory leg runs and how many fail. A smaller
failure count is the interesting direction: it means one now passes, which is the trigger the KDoc
already names for deleting the annotation. The report also prints #81's
grepcount beside thebaseline, so a stale baseline shows up on the run rather than only when the emulator disagrees.
The gating legs get the shape without the comparison — they run the whole suite, so comparing there
would announce a deviation five times a run.
Mutation, both directions
1. A fourth failing
@FailsOnEmulatorApi37test — run on real CI, from a throwaway branch(#112, closed and deleted). Run
32899133552, advisory job97968820644, verbatim:Those three landed as real
noticeannotations on the check run, and the report names the newtest, which is the thing a count alone cannot do. The advisory job's conclusion was
failure— the conclusion it has on every PR. It stayed
continue-on-error: trueand stayed out of therequired contexts.
That run's run-level conclusion was also red, and that is the probe rather than the report:
the probe fails unconditionally, and
@FailsOnEmulatorApi37only removes it from the API 37gating leg, so legs 33–36 ran it. API 33's own report says so —
Starting 60 tests,There was 1 failure(s), the one failure beingMutationProbeApi37Test. The API 37 gating legstayed green.
2. The baseline number changed alone — run locally against the captured stdout of a real CI
run,
32865281555, because the API 37 emulator image cannot be run on this host (CLAUDE.md:local emulators are API 33–36).
sedon the committed const,3→4, nothing else touched:Restored to
3immediately after; the same log against the committed baseline printsbaseline: matches (3 expected, 3 failed)and no notice.3. An unreadable baseline is a deviation too, added after review pointed out that the
^-anchoredsedmeans indenting the const into an object would silently switch the comparisonoff while the table kept printing — the ticket's own failure mode wearing a different hat.
Verified locally three ways against the same real log: file absent → announces; const indented
into an
object→ announces; committed file unchanged → silent.Degenerate shapes
A diagnostic that crashes, or that invents a number, is worse than none. Checked against real and
constructed logs: no log file, empty log file,
Starting 0 tests(the framework-restart shape),and a stale
app/buildXML with no test run in the log. All exit 0, none report a bogus0 vs 3,and the XML is read only when the runner said a run started — which matters on
tools/local-emulator/run-e2e.sh, where one checkout drives several API levels.Not covered
The one thing I could not do on this host is run the API 37 leg locally, so mutation 2 and the
degenerate shapes were exercised against captured CI output rather than a live emulator.
Mutation 1 and the abort itself were seen on real CI, in run
32899133552(
Expected 4 tests, received 3…INSTRUMENTATION_ABORTED), so the truncation capture is notinferred.
The negative control, on this branch's own run
Run
32918773988, advisory job98028007674— the same code, nothing mutated:Zero
noticeannotations on that job; the mutation run had three. That pair is thedemonstration — silent when it matches, loud when it does not. The last line is there because
GitHub's check-run API returns
summary: nullfor a job summary, so a write that silently did nothappen would be invisible; it confirms
GITHUB_STEP_SUMMARYreaches the emulator-runner's childprocess.
E2E API 33was red on that run with one failure,ReattachOnLaunchTest.doesNotOverwriteAPickTheUserHasAlreadyMade— that is #49, a known flake, andthe report's own
failed tests:list is how it was identified, without opening a log. Re-runningthe failed legs put the run at conclusion
success:E2E API 33green, all eight requiredcontexts green, and the advisory job on that attempt (
98029549102) reportingexpected 3 / received 3 / failed 3 / completed cleanly: no,baseline: matches, zero notices.The run's conclusion is
successwhile the advisory job's conclusion isfailure. That is theconstraint holding, demonstrated rather than asserted: the report ran, compared, found a match, and
the advisory red still did not gate anything.