Stop the cancel tests losing their race on a loaded runner #240

Merged
JMR-dev merged 1 commits from fix/cancel-tests-need-a-slower-encode into main 2026-09-06 06:20:58 +00:00
JMR-dev commented 2026-09-06 06:04:06 +00:00 (Migrated from github.com)

Fixes a flake I introduced in ad2a75d (#236) and d293646 (#237).

Both cancellation tests wait for SessionState.RUNNING and then cancel. That is not enough. The conversion one passed four consecutive local runs and all five CI legs, then failed the API 34 and 35 legs of the next PR with state=COMPLETED rc=0, on a diff that could not reach it.

On a loaded runner the thread that observed RUNNING can be descheduled long enough for a short encode to finish before it calls cancel. A longer timeout does not help — the wait already succeeded.

Two changes, because neither is sufficient alone

A slower encode. The conversion test now targets WEBM_VP9 at BEST, the slowest thing FFmpegCommandBuilder emits: libvpx-vp9 -crf 31 -b:v 0, with -deadline realtime added only on FAST. Probed on an API 34 emulator:

target still RUNNING at 1 s? finished by
MP4_H265 / BEST (x265 -preset medium) no well under 1 s
WEBM_VP9 / BEST yes 2 s

A bounded retry. An attempt whose session finished before the cancel landed has not tested anything — that is a miss, not a failure, so it is retried. Only exhausting five attempts fails, and the message reports every attempt's state and return code, so a real breakage is distinguishable from a slow machine.

The retry does not soften the test

With FFmpegKit.cancel removed from both engines, every attempt ends COMPLETED, so both tests still fail:

AssertionError: never interrupted a running session in 5 attempts, so either every encode
  finished first or cancellation does not reach it:
  [state=COMPLETED rc=0, state=COMPLETED rc=0, state=COMPLETED rc=0, state=COMPLETED rc=0, state=COMPLETED rc=0]

Two clean runs beforehand at 64/0/0/3.

The join test gets the same treatment. It has not flaked yet, but it is the same mechanism and the same fragility, and finding out on CI again is not worth the round trip.

🤖 Generated with Claude Code

Fixes a flake I introduced in `ad2a75d` (#236) and `d293646` (#237). Both cancellation tests wait for `SessionState.RUNNING` and then cancel. **That is not enough.** The conversion one passed four consecutive local runs and all five CI legs, then failed the **API 34 and 35** legs of the next PR with `state=COMPLETED rc=0`, on a diff that could not reach it. On a loaded runner the thread that observed `RUNNING` can be descheduled long enough for a short encode to finish before it calls `cancel`. A longer timeout does not help — the wait already succeeded. ## Two changes, because neither is sufficient alone **A slower encode.** The conversion test now targets `WEBM_VP9` at `BEST`, the slowest thing `FFmpegCommandBuilder` emits: `libvpx-vp9 -crf 31 -b:v 0`, with `-deadline realtime` added **only** on `FAST`. Probed on an API 34 emulator: | target | still `RUNNING` at 1 s? | finished by | |---|---|---| | `MP4_H265` / `BEST` (x265 `-preset medium`) | no | well under 1 s | | `WEBM_VP9` / `BEST` | **yes** | 2 s | **A bounded retry.** An attempt whose session finished before the cancel landed has not tested anything — that is a *miss*, not a failure, so it is retried. Only exhausting five attempts fails, and the message reports every attempt's state and return code, so a real breakage is distinguishable from a slow machine. ## The retry does not soften the test With `FFmpegKit.cancel` removed from **both** engines, every attempt ends `COMPLETED`, so both tests still fail: ``` AssertionError: never interrupted a running session in 5 attempts, so either every encode finished first or cancellation does not reach it: [state=COMPLETED rc=0, state=COMPLETED rc=0, state=COMPLETED rc=0, state=COMPLETED rc=0, state=COMPLETED rc=0] ``` Two clean runs beforehand at `64/0/0/3`. The join test gets the same treatment. It has not flaked yet, but it is the same mechanism and the same fragility, and finding out on CI again is not worth the round trip. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign in to join this conversation.