Fixes a flake I introduced in ad2a75d (#236) and d293646 (#237).
Both cancellation tests wait for SessionState.RUNNING and then cancel. That is not enough. The conversion one passed four consecutive local runs and all five CI legs, then failed the API 34 and 35 legs of the next PR with state=COMPLETED rc=0, on a diff that could not reach it.
On a loaded runner the thread that observed RUNNING can be descheduled long enough for a short encode to finish before it calls cancel. A longer timeout does not help — the wait already succeeded.
Two changes, because neither is sufficient alone
A slower encode. The conversion test now targets WEBM_VP9 at BEST, the slowest thing FFmpegCommandBuilder emits: libvpx-vp9 -crf 31 -b:v 0, with -deadline realtime added only on FAST. Probed on an API 34 emulator:
target
still RUNNING at 1 s?
finished by
MP4_H265 / BEST (x265 -preset medium)
no
well under 1 s
WEBM_VP9 / BEST
yes
2 s
A bounded retry. An attempt whose session finished before the cancel landed has not tested anything — that is a miss, not a failure, so it is retried. Only exhausting five attempts fails, and the message reports every attempt's state and return code, so a real breakage is distinguishable from a slow machine.
The retry does not soften the test
With FFmpegKit.cancel removed from both engines, every attempt ends COMPLETED, so both tests still fail:
AssertionError: never interrupted a running session in 5 attempts, so either every encode
finished first or cancellation does not reach it:
[state=COMPLETED rc=0, state=COMPLETED rc=0, state=COMPLETED rc=0, state=COMPLETED rc=0, state=COMPLETED rc=0]
Two clean runs beforehand at 64/0/0/3.
The join test gets the same treatment. It has not flaked yet, but it is the same mechanism and the same fragility, and finding out on CI again is not worth the round trip.
Fixes a flake I introduced in `ad2a75d` (#236) and `d293646` (#237).
Both cancellation tests wait for `SessionState.RUNNING` and then cancel. **That is not enough.** The conversion one passed four consecutive local runs and all five CI legs, then failed the **API 34 and 35** legs of the next PR with `state=COMPLETED rc=0`, on a diff that could not reach it.
On a loaded runner the thread that observed `RUNNING` can be descheduled long enough for a short encode to finish before it calls `cancel`. A longer timeout does not help — the wait already succeeded.
## Two changes, because neither is sufficient alone
**A slower encode.** The conversion test now targets `WEBM_VP9` at `BEST`, the slowest thing `FFmpegCommandBuilder` emits: `libvpx-vp9 -crf 31 -b:v 0`, with `-deadline realtime` added **only** on `FAST`. Probed on an API 34 emulator:
| target | still `RUNNING` at 1 s? | finished by |
|---|---|---|
| `MP4_H265` / `BEST` (x265 `-preset medium`) | no | well under 1 s |
| `WEBM_VP9` / `BEST` | **yes** | 2 s |
**A bounded retry.** An attempt whose session finished before the cancel landed has not tested anything — that is a *miss*, not a failure, so it is retried. Only exhausting five attempts fails, and the message reports every attempt's state and return code, so a real breakage is distinguishable from a slow machine.
## The retry does not soften the test
With `FFmpegKit.cancel` removed from **both** engines, every attempt ends `COMPLETED`, so both tests still fail:
```
AssertionError: never interrupted a running session in 5 attempts, so either every encode
finished first or cancellation does not reach it:
[state=COMPLETED rc=0, state=COMPLETED rc=0, state=COMPLETED rc=0, state=COMPLETED rc=0, state=COMPLETED rc=0]
```
Two clean runs beforehand at `64/0/0/3`.
The join test gets the same treatment. It has not flaked yet, but it is the same mechanism and the same fragility, and finding out on CI again is not worth the round trip.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Fixes a flake I introduced in
ad2a75d(#236) andd293646(#237).Both cancellation tests wait for
SessionState.RUNNINGand then cancel. That is not enough. The conversion one passed four consecutive local runs and all five CI legs, then failed the API 34 and 35 legs of the next PR withstate=COMPLETED rc=0, on a diff that could not reach it.On a loaded runner the thread that observed
RUNNINGcan be descheduled long enough for a short encode to finish before it callscancel. A longer timeout does not help — the wait already succeeded.Two changes, because neither is sufficient alone
A slower encode. The conversion test now targets
WEBM_VP9atBEST, the slowest thingFFmpegCommandBuilderemits:libvpx-vp9 -crf 31 -b:v 0, with-deadline realtimeadded only onFAST. Probed on an API 34 emulator:RUNNINGat 1 s?MP4_H265/BEST(x265-preset medium)WEBM_VP9/BESTA bounded retry. An attempt whose session finished before the cancel landed has not tested anything — that is a miss, not a failure, so it is retried. Only exhausting five attempts fails, and the message reports every attempt's state and return code, so a real breakage is distinguishable from a slow machine.
The retry does not soften the test
With
FFmpegKit.cancelremoved from both engines, every attempt endsCOMPLETED, so both tests still fail:Two clean runs beforehand at
64/0/0/3.The join test gets the same treatment. It has not flaked yet, but it is the same mechanism and the same fragility, and finding out on CI again is not worth the round trip.
🤖 Generated with Claude Code