The advisory baseline check has announced a deviation on every PR since #113:
the tree "carries 4 tests marked @FailsOnEmulatorApi37" where it carries three
and FAILS_ON_EMULATOR_API37_BASELINE says three. The fourth is a KDoc in
Media3EngineTest saying the opposite -- "Deliberately not
`@FailsOnEmulatorApi37`: nothing here decodes or encodes" -- which the old
matcher counted because it looked for the string anywhere on any line.
Neither ingredient was wrong on its own, and the number is not the real damage.
#83 added this check so that a new failure joining the known ones could not be
invisible; a notice that is wrong every single time teaches everyone to skim
past deviation notices, which is precisely the signal it was built to create.
Editing the baseline to 4 would have silenced it by breaking it -- the check
would then have been wrong the moment someone added or removed a real marker.
Anchor the pattern at line start and require whitespace or end-of-line after the
name. The second half is the part that is easy to get wrong: "only the
annotation on a line of its own" also stops counting `@FailsOnEmulatorApi37
@Test`, which is legal Kotlin, and undercounting is the dangerous direction --
it hides a genuine new marker, the one thing this exists to catch. Measured
against a fixture carrying every shape at once: the old matcher 5, own-line-only
2, this one 3; on the real tree 4 / 3 / 3, so the baseline is untouched.
`grep -v import` goes too, since `^[[:space:]]*@` cannot match an import.
The check is a pure function of the working tree, so the fixture is committed
and e2e-report-shape-test.sh runs the real report against it -- inside a
throwaway repo root, which the script finds from BASH_SOURCE, so no knob had to
be added that could point the live count somewhere else. The fixture sits under
.github/, where Gradle does not compile it and :app's ktlint and detekt do not
see it; running the report against the real root with it committed still
reports 3.
Every other path through the report is byte-identical to the previous version on
both stdout and the job summary -- passing, failing, wedged, no-run, and
advisory-with-an-unreadable-baseline all diff empty -- and the two advisory legs
differ only by the false line disappearing. No job's status or pass/fail rules
change; the advisory leg stays continue-on-error and stays red by design.
The test is deliberately not wired into CI: adding a step to Static analysis
would add a new way for a gating job to go red, which #120 ruled out. shellcheck
still covers the file, since that step reads `git ls-files '*.sh'`.
Closes#120
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The report added by #111 runs on every path out of e2e-run.sh, including the
wedge, and until now it answered a question it had not been asked. On job
98035980326 -- API 34, a docs-only PR -- it printed `received: 59` and
`completed cleanly: yes` six seconds before `##[warning] ... WEDGED`, for a leg
the WEDGE_TIMEOUT had killed 22 minutes in. `completed cleanly` means only
"instrumentation was not aborted", which was true; a reader scanning the table
had to notice a separate warning line to learn the leg had died.
The wedge cannot be read out of the log, which is why it is passed in: a wedge
is gradle never returning, so gradle printed no verdict, no truncation line and
no INSTRUMENTATION_ABORTED, and the log it leaves is the log of a run that just
stops. Only e2e-run.sh saw `timeout` exit 124. It now derives that fact once and
tells the report as E2E_WEDGED_AFTER, and reuses the same variable for
capture_wedge so the two cannot drift.
The table gains a `wedged:` row above `completed cleanly`, and `completed
cleanly` flips to no -- but only where it would have said yes. An abort already
says no and names the abort, which the wedge row does not, and a run that left
no evidence still says unknown; a wedge on top of either prints both facts.
`received`'s source line told the same lie in the same table -- "the run was not
truncated, so every expected test reported" is only "gradle never got as far as
saying so" when the leg was killed -- so it is qualified on that path. The
number itself is unchanged, and so is `failed: unknown`: gradle printed no
summary line, so that count genuinely is not knowable.
Nothing here decides anything. No exit status, no pass/fail rule, no baseline
comparison and no `::notice::` behaviour changes; the leg already failed
correctly and still does.
Verified against captured CI output rather than a live emulator, as #111 was and
for the same reason -- this host cannot run API 37 and cannot wedge on demand.
Four real logs (the wedged leg, a green API 34 leg, a failing gating leg, and an
advisory leg with its baseline deviation) through both versions of the script,
in both env states, comparing stdout and the job summary: only the wedged run
with the signal set differs, byte for byte.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The job summary is the deliverable #83 asked for -- "readable without opening a
log" -- and GitHub exposes no API that reads a job summary back: the check-run
output for the advisory job returns summary: null, so a write that silently did
not happen would be invisible to everything except a human on the run page. The
step log can be read, so it now carries one line saying which of the two
happened, including the case where GITHUB_STEP_SUMMARY is unset entirely, which
is what running the script by hand looks like.
"A comparison was asked for" and "a number was found to compare against" were one
variable, and collapsing them put the report one refactor away from being the
thing #83 filed. The sed that reads FAILS_ON_EMULATOR_API37_BASELINE is anchored
at the line start, so indenting the const into an object -- or renaming it, or
moving it -- empties it, and the old code then skipped the whole comparison while
the table kept printing exactly as before. Silent, and indistinguishable from a
run that matched.
Now an unreadable baseline is itself a deviation, with the notice naming the
const so the fix is obvious. Verified against the real captured log of run
32865281555 three ways: baseline file absent, const indented into an object, and
the committed file unchanged -- the first two announce, the third stays silent.
That job is continue-on-error and red on every PR by design, which CLAUDE.md
states plainly -- and that instruction is exactly why nobody reads it. Nothing in
a red X separates "the known three" from "the known three plus yours".
A bare failure count would not have fixed it, and this is measured rather than
assumed. The run is usually truncated: seven of eight advisory runs read on
2026-08-25 ended in `Test run failed to complete. Expected 3 tests, received 2.`
with INSTRUMENTATION_ABORTED, and one did not. A count taken from a truncated run
misleads in both directions -- a fourth marked test can still yield the same
number if the abort lands earlier, and the known set getting worse can lower it.
The test XML does not rescue it either, which was the thing worth checking before
building on it: it IS written for an aborted run, and it reports a tidy
tests="3" failures="3" for a run the runner had just described as truncated. So
the XML is the authority on how many results landed, the runner's own output is
the only authority on whether the run finished, and the report reads both and
says which number came from where.
The baseline is one number beside the marker, because the marker means "cannot
pass on this image": the count is both how many tests the advisory leg runs and
how many should fail. A smaller failure count is the interesting direction -- it
means one now passes, which is the documented trigger for deleting the
annotation.
Nothing about the job's status changes. It stays continue-on-error, stays red,
stays out of the required contexts; a deviation is a ::notice::, never an
::error::. The report is a separate script so it can be run against a real log
saved from a real CI run, which is how the comparison was shown to fire.
The gating legs get the shape without the comparison: they run the whole suite,
so comparing there would announce a deviation five times a run -- but a truncated
run reporting fewer results than it ran is what #108 looks like, and "completed
cleanly" is the field that would show it.
Closes#83