Three defects found by truncating a real download and killing one mid-flight.
Matroska and WebM write the duration into the header, so ffprobe reports the
full 212.8s of a video.mkv cut off after a kilobyte -- exit 0, plausible
answer. verify_video therefore returned OK, the corrupt file was never
re-downloaded, ffmpeg extracted the 0.02s of audio it could find, and the run
reported "1 ok" with an empty transcript. The size recorded at download time
is the only evidence the bytes are still there, so it is now checked whenever
it is known rather than only as a fallback when the duration is unreadable.
The audio stage now takes the expected duration and rejects an extraction that
does not match it. ffmpeg exits 0 on a truncated container, so without this a
damaged source yields a confident transcript of near-silence, which is a worse
outcome than a failed job. It compares against the video's own probed duration
rather than state.duration_s, which can come from playlist metadata.
aria2 saves its control file every 60s by default. Since that file is what a
resume reads, a kill -9 inside the first minute preserved a control file
recording zero completed pieces: measured 0/13 on the sample, so the "resume"
re-downloaded the lot while reporting a partial. At --auto-save-interval=20 the
same kill preserves 2/13 pieces and the resumed download is byte-identical to a
clean one.
The duration tolerance moves to media/ffmpeg.py, which both callers already
import, instead of being restated per call site.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>