docs: record what a resume actually preserves, and date the benchmark

The resume section implied a .part file's size reflected progress; with falloc
it is full-size from the first second and only the checkpointed pieces are
skipped. Adds the corrupt-video and truncation rules, and the new
--auto-save-interval row.

README-intel.md's 5.2x figure does not reproduce on this box today (~4.0x from
the repo's own bench.py, GPU/CPU ratio unchanged), so the table is marked as a
point-in-time measurement rather than a target to chase.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-09-13 16:19:00 -05:00
co-authored by Claude Opus 5
parent 3bff193f92
commit 4d19d62077
2 changed files with 19 additions and 3 deletions
+5
View File
@@ -129,6 +129,11 @@ Measured on Iris Xe (80 EU) + i5-1145G7, `large-v3-turbo-int8`, 121 s of audio:
| GPU | ~23 s | ~5.2x | ~24 s | 1.0 |
| CPU | ~39 s | ~3.1x | ~147 s | 3.8 |
> These are a point-in-time measurement. Re-running `scripts/bench.py GPU sample.wav`
> on the same box in September 2026 gave 29-32 s (~4.0x) rather than 23 s, with the
> GPU/CPU ratio unchanged. Treat the ratio as the durable result and re-measure the
> absolute numbers on your own machine before calling a slower run a regression.
The iGPU is ~1.7x faster and uses ~6x less CPU time. Note it still occupies **one
full core** — the plugin busy-waits — so the offload is 3.8 cores down to 1.0, not
to zero. Budget for that if transcription runs alongside other CPU work.
+14 -3
View File
@@ -47,9 +47,19 @@ stage whose output does not check out.
| ok | — | — | nothing |
Killing a run mid-download resumes rather than restarting, because yt-dlp keeps
its `.part` and aria2 keeps its control file. A `.part` **without** that control
file is treated as unusable and restarted: aria2 writes segments out of order, so
such a file is sparse with holes rather than a valid prefix.
its `.part` and aria2 keeps its control file. How much survives depends on the
checkpoint interval, not on the size of the `.part`: with `--file-allocation=falloc`
the file is full-size from the first second, and only the pieces recorded in the
control file are skipped. A `.part` **without** that control file is discarded and
restarted, because aria2 writes segments out of order, so such a file is sparse
with holes rather than a valid prefix -- resuming from its length would produce a
video that stops being the source partway through.
A video that exists but fails verification is unlinked before re-downloading.
yt-dlp's overwrite default is off, so it would otherwise report the corrupt file
as "already downloaded" and skip it forever. Verification checks the recorded byte
count as well as the duration: Matroska and WebM write the duration into the
header, so a file truncated to a kilobyte still probes as its full length.
Deleting one subtitle file re-renders it from the segments saved in
`transcript.json` instead of re-transcribing the whole recording.
@@ -110,6 +120,7 @@ this adds is the flags that do:
| `--lowest-speed-limit=50K` | 0 (off) | without it a wedged connection hangs forever |
| `--disk-cache=64M` | 16M | sixteen writers at out-of-order offsets thrash a small cache |
| `--file-allocation=falloc` | none | instant on btrfs/ext4/xfs; avoids fragmenting multi-GB files |
| `--auto-save-interval=20` | 60 | the control file is what a resume reads; a kill inside the first minute otherwise preserves zero completed pieces |
| `--log=logs/aria2.log` | none | the console level hides everything, leaving no forensics |
`--continue` is deliberately never passed — see the `.part` note above.