Report the advisory API 37 job's expected/received/failed counts, so a new failure is not invisible #83

Closed
opened 2026-08-25 03:00:04 +00:00 by JMR-dev · 0 comments
JMR-dev commented 2026-08-25 03:00:04 +00:00 (Migrated from github.com)

Requested 2026-08-25. The advisory API 37 job is red on every PR by design, so a new failure joining it is invisible — nothing distinguishes "the known ones" from "the known ones plus yours". Flagged during the #56 review and never closed.

Constraint, stated first

Nothing about the job's status or pass/fail rules changes. It stays continue-on-error: true, it
stays red, it stays out of ruleset 21117412's eight required contexts. This ticket adds a
report at the end of the job and nothing else. If the implementation makes the job able to fail
the build, it is wrong.

A bare failure count is the wrong signal here, and this is measured

From run 32800638011, the advisory job's own log:

Execute org.libremediaconverter.convert.Media3EngineTest.transcodesH264ToH265AndReportsProgress: FAILED
Execute org.libremediaconverter.saf.SafPickerRoundTripTest.thePickedInputSurvivesARealRotation: FAILED
Test run failed to complete. Expected 3 tests, received 2.
  onError: commandError=false message=INSTRUMENTATION_ABORTED: System has crashed.

Three tests carry the marker; only two ever report. thePickedInputSurvivesARealRotation takes
the framework down, the instrumentation aborts, and Media3EngineTest.runsFromAThreadWithNoLooper
never runs at all. So "failures = 2" is today's number, and it is 2 because a crash truncated the
run
, not because one of the three passed.

That makes a naive failure count actively misleading in both directions:

  • A fourth test joining the marker could still yield "2 failures" if the abort happens earlier.
  • The known set regressing (say the abort moving earlier still) would lower the count.

What to record instead

The runner already prints the honest triple. Capture all of it:

field today
expected 3
received 2
failed 2
completed cleanly no — INSTRUMENTATION_ABORTED

Then compare against a committed baseline and say plainly whether it matches. Sketch, not a
prescription:

  • Write the four values to $GITHUB_STEP_SUMMARY so they are readable without opening a log, and
    emit a ::notice:: (never ::error:: — see the constraint) when they differ from the baseline.
  • Keep the baseline next to the marker's definition so the two move together. Whoever adds a test to
    @FailsOnEmulatorApi37 should have to update one number in one place, and the diff should say so.
  • Check whether a test XML is even produced when the run aborts before building on it. If it is
    absent or partial, parse the runner's own "Expected N tests, received M" line instead; do not
    assume the XML exists.

Why this is worth doing

Three separate subagents have flagged this job's by-design failure as possibly their own change
breaking something. CLAUDE.md now says plainly that it is red on every PR — but that instruction
also means nobody looks, which is exactly the condition under which a real regression hides.
A green run is not evidence those three pass; today, neither is a red one evidence that only those
three failed.

Done means

  • The job prints expected/received/failed and whether the run completed, at the end, every time.
  • A deviation from the committed baseline is visibly announced without changing the job's conclusion.
  • Mutation: add a fourth test carrying @FailsOnEmulatorApi37 that fails, confirm the report
    says so and the job's conclusion is unchanged; remove it. Then change the baseline number alone and
    confirm the report notices that too — a comparison that cannot be shown to fire is not a comparison.
  • The related count check from #81 still holds:
    grep -rn "@FailsOnEmulatorApi37" app/src/androidTest --include='*.kt' | grep -v import | grep -c FailsOn
_Requested 2026-08-25. The advisory API 37 job is red on every PR by design, so a **new** failure joining it is invisible — nothing distinguishes "the known ones" from "the known ones plus yours". Flagged during the #56 review and never closed._ ### Constraint, stated first **Nothing about the job's status or pass/fail rules changes.** It stays `continue-on-error: true`, it stays red, it stays out of ruleset `21117412`'s eight required contexts. This ticket adds a **report** at the end of the job and nothing else. If the implementation makes the job able to fail the build, it is wrong. ### A bare failure count is the wrong signal here, and this is measured From run `32800638011`, the advisory job's own log: ``` Execute org.libremediaconverter.convert.Media3EngineTest.transcodesH264ToH265AndReportsProgress: FAILED Execute org.libremediaconverter.saf.SafPickerRoundTripTest.thePickedInputSurvivesARealRotation: FAILED Test run failed to complete. Expected 3 tests, received 2. onError: commandError=false message=INSTRUMENTATION_ABORTED: System has crashed. ``` **Three tests carry the marker; only two ever report.** `thePickedInputSurvivesARealRotation` takes the framework down, the instrumentation aborts, and `Media3EngineTest.runsFromAThreadWithNoLooper` never runs at all. So "failures = 2" is today's number, and it is 2 **because a crash truncated the run**, not because one of the three passed. That makes a naive failure count actively misleading in both directions: - A **fourth** test joining the marker could still yield "2 failures" if the abort happens earlier. - The known set **regressing** (say the abort moving earlier still) would *lower* the count. ### What to record instead The runner already prints the honest triple. Capture all of it: | field | today | |---|---| | expected | 3 | | received | 2 | | failed | 2 | | completed cleanly | **no** — `INSTRUMENTATION_ABORTED` | Then compare against a **committed baseline** and say plainly whether it matches. Sketch, not a prescription: - Write the four values to `$GITHUB_STEP_SUMMARY` so they are readable without opening a log, and emit a `::notice::` (never `::error::` — see the constraint) when they differ from the baseline. - Keep the baseline next to the marker's definition so the two move together. Whoever adds a test to `@FailsOnEmulatorApi37` should have to update one number in one place, and the diff should say so. - **Check whether a test XML is even produced when the run aborts** before building on it. If it is absent or partial, parse the runner's own "Expected N tests, received M" line instead; do not assume the XML exists. ### Why this is worth doing Three separate subagents have flagged this job's by-design failure as possibly their own change breaking something. `CLAUDE.md` now says plainly that it is red on every PR — but that instruction also means **nobody looks**, which is exactly the condition under which a real regression hides. A green run is not evidence those three pass; today, neither is a red one evidence that only those three failed. ### Done means - The job prints expected/received/failed and whether the run completed, at the end, every time. - A deviation from the committed baseline is visibly announced without changing the job's conclusion. - **Mutation:** add a fourth test carrying `@FailsOnEmulatorApi37` that fails, confirm the report says so and the job's conclusion is unchanged; remove it. Then change the baseline number alone and confirm the report notices that too — a comparison that cannot be shown to fire is not a comparison. - The related count check from #81 still holds: `grep -rn "@FailsOnEmulatorApi37" app/src/androidTest --include='*.kt' | grep -v import | grep -c FailsOn`
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: JMR-dev/LibreMediaConverter#83