From 0702916229a45fc9f20a70b05bc82653b4964d5a Mon Sep 17 00:00:00 2001 From: Jason Ross Date: Tue, 25 Aug 2026 16:04:13 -0500 Subject: [PATCH 1/3] Say what the advisory API 37 job actually found, so a new failure is not invisible That job is continue-on-error and red on every PR by design, which CLAUDE.md states plainly -- and that instruction is exactly why nobody reads it. Nothing in a red X separates "the known three" from "the known three plus yours". A bare failure count would not have fixed it, and this is measured rather than assumed. The run is usually truncated: seven of eight advisory runs read on 2026-08-25 ended in `Test run failed to complete. Expected 3 tests, received 2.` with INSTRUMENTATION_ABORTED, and one did not. A count taken from a truncated run misleads in both directions -- a fourth marked test can still yield the same number if the abort lands earlier, and the known set getting worse can lower it. The test XML does not rescue it either, which was the thing worth checking before building on it: it IS written for an aborted run, and it reports a tidy tests="3" failures="3" for a run the runner had just described as truncated. So the XML is the authority on how many results landed, the runner's own output is the only authority on whether the run finished, and the report reads both and says which number came from where. The baseline is one number beside the marker, because the marker means "cannot pass on this image": the count is both how many tests the advisory leg runs and how many should fail. A smaller failure count is the interesting direction -- it means one now passes, which is the documented trigger for deleting the annotation. Nothing about the job's status changes. It stays continue-on-error, stays red, stays out of the required contexts; a deviation is a ::notice::, never an ::error::. The report is a separate script so it can be run against a real log saved from a real CI run, which is how the comparison was shown to fire. The gating legs get the shape without the comparison: they run the whole suite, so comparing there would announce a deviation five times a run -- but a truncated run reporting fewer results than it ran is what #108 looks like, and "completed cleanly" is the field that would show it. Closes #83 --- .github/scripts/e2e-report-shape.sh | 270 ++++++++++++++++++ .github/scripts/e2e-run.sh | 37 ++- .github/workflows/status_check.yml | 8 + CLAUDE.md | 9 + .../FailsOnEmulatorApi37.kt | 29 ++ 5 files changed, 352 insertions(+), 1 deletion(-) create mode 100755 .github/scripts/e2e-report-shape.sh diff --git a/.github/scripts/e2e-report-shape.sh b/.github/scripts/e2e-report-shape.sh new file mode 100755 index 0000000..865bef7 --- /dev/null +++ b/.github/scripts/e2e-report-shape.sh @@ -0,0 +1,270 @@ +#!/usr/bin/env bash +# +# Reports the SHAPE of an instrumented run -- how many tests were expected, how many +# reported, how many failed, and whether the run completed at all -- to the step log and to +# the job summary. In advisory mode it also compares that shape against a committed baseline +# and says plainly whether it matches. +# +# WHY THIS EXISTS (#83): the advisory API 37 leg is red on every PR by design, so a NEW failure +# joining the known ones is invisible -- nothing in a red X distinguishes "the known ones" from +# "the known ones plus yours". CLAUDE.md tells everyone not to read that job's red as their +# change breaking something, which is correct, and which also means nobody looks. +# +# WHY NOT A BARE FAILURE COUNT, measured rather than assumed. On this image the run is usually +# truncated: `Test run failed to complete. Expected 3 tests, received 2.` with +# `INSTRUMENTATION_ABORTED: System has crashed.` A count taken from a truncated run misleads in +# both directions -- a fourth marked test can still yield the same number if the abort lands +# earlier, and the known set getting worse can LOWER it. So all four fields are recorded, and +# the one saying the run was truncated is recorded with them. +# +# WHY IT IS A SEPARATE SCRIPT rather than a function inside e2e-run.sh: it is a pure seam. It +# reads a captured log plus the test XML and writes a report, so it can be run against a REAL +# log saved from a REAL CI run -- which is how the baseline comparison was shown to fire +# without waiting on an emulator. `git ls-files '*.sh'` also picks it up for shellcheck for +# free. +# +# THIS SCRIPT NEVER FAILS A RUN. It is a diagnostic, and e2e-run.sh's header explains why that +# rule is absolute here. Every field defaults to `unknown` and every comparison is guarded, +# because an unset variable under `set -u`, or a `[ "" -eq 3 ]`, is exactly how a diagnostic +# becomes the thing that turns a leg red. It exits 0 unconditionally. +# +# Usage: +# e2e-report-shape.sh