JMR-devandClaude Opus 5 3bff193f92 fix(verify): catch a truncated container, and checkpoint aria2 often enough to resume
Three defects found by truncating a real download and killing one mid-flight.

Matroska and WebM write the duration into the header, so ffprobe reports the
full 212.8s of a video.mkv cut off after a kilobyte -- exit 0, plausible
answer. verify_video therefore returned OK, the corrupt file was never
re-downloaded, ffmpeg extracted the 0.02s of audio it could find, and the run
reported "1 ok" with an empty transcript. The size recorded at download time
is the only evidence the bytes are still there, so it is now checked whenever
it is known rather than only as a fallback when the duration is unreadable.

The audio stage now takes the expected duration and rejects an extraction that
does not match it. ffmpeg exits 0 on a truncated container, so without this a
damaged source yields a confident transcript of near-silence, which is a worse
outcome than a failed job. It compares against the video's own probed duration
rather than state.duration_s, which can come from playlist metadata.

aria2 saves its control file every 60s by default. Since that file is what a
resume reads, a kill -9 inside the first minute preserved a control file
recording zero completed pieces: measured 0/13 on the sample, so the "resume"
re-downloaded the lot while reporting a partial. At --auto-save-interval=20 the
same kill preserves 2/13 pieces and the resumed download is byte-identical to a
clean one.

The duration tolerance moves to media/ffmpeg.py, which both callers already
import, instead of being restated per call site.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 16:17:14 -05:00

ccn-transcribe

Download a video, extract its audio losslessly as FLAC, and transcribe it with Whisper on whatever GPU the machine actually has.

Resumable from any completed stage, with errors that tell you what to do next.

uv sync
uv run ccn-transcribe doctor                      # check the environment first
uv run ccn-transcribe run 'https://youtu.be/...'
uv run ccn-transcribe run -i urls.txt --formats txt,srt

Hardware setup for the Intel path is documented separately in README-intel.md.

What it does

yt-dlp (aria2c)  ->  ffmpeg  ->  Whisper        ->  txt / srt / vtt / json
   video.mkv         audio.flac   OpenVINO GenAI

Each video becomes one job directory under --workdir (default transcripts/):

transcripts/jobs/youtube-<id>/
  state.json          media/video.mkv     out/transcript.{txt,srt,vtt,json}
  events.jsonl        media/audio.flac    logs/{job,aria2}.log

Resume

Artifacts on disk are the source of truth. state.json records only what the filesystem cannot answer — chiefly why a file is absent. There is deliberately no stored "current stage": persisting one is how "marked done but the file is gone" bugs happen. Every run verifies what is actually present (ffprobe duration > 0, non-empty, audio length matching the video) and enters at the first stage whose output does not check out.

Transcript FLAC Video Runs
missing ok anything transcribe, outputs
missing missing ok extract, transcribe, outputs
missing missing missing everything
ok — — nothing

Killing a run mid-download resumes rather than restarting, because yt-dlp keeps its .part and aria2 keeps its control file. A .part without that control file is treated as unusable and restarted: aria2 writes segments out of order, so such a file is sparse with holes rather than a valid prefix.

Deleting one subtitle file re-renders it from the segments saved in transcript.json instead of re-transcribing the whole recording.

--dry-run prints the plan for every job without touching anything.

Keeping or discarding media

Video and audio are kept by default.

Flag Effect
--no-retain-video delete the video once the FLAC is extracted and verified
--no-retain-audio delete the FLAC once every output is written
--no-retain both

Deletion always happens after the next stage is committed, so interrupting a run is still safe and resume still works. What you actually give up is re-run flexibility: re-transcribing later with a different --language, --model or --task will need another download.

These flags prompt once per run. On a non-TTY they refuse outright rather than proceeding — a piped or cron invocation should never delete gigabytes with nobody seeing the warning. Pass --yes to confirm non-interactively.

Hardware routing

Detection runs in two phases: a cheap hardware probe that imports nothing, then a toolchain probe that may import. The toolchain probe is only reached once the hardware is confirmed, which is what keeps torch off a machine with no NVIDIA or AMD GPU.

Order is NVIDIA → AMD → Intel GPU → CPU, and --device overrides it.

Fallback happens at runtime, not at detection: constructing the backend and its first generate() are both guarded, because a /dev/dri permission failure and an OpenCL JIT failure surface there rather than at device enumeration. A backend that fails is demoted for the rest of the process, so a long batch does not retry it once per job. An explicit --device never falls back silently — getting the CPU when you asked for the GPU would make any measurement a lie.

NVIDIA and AMD are interface-only in this build. Detection and routing are real and the error says what to implement, but there is no inference code: the cached model is OpenVINO IR, which cannot load on CUDA or ROCm, and neither path could be tested here. CTranslate2 has no ROCm support, so they are genuinely two implementations rather than one.

Download speed

yt-dlp already passes aria2c -x16 -j16 -s16 --min-split-size 1M, and -x16 is aria2's per-server cap, so re-sending connection counts changes nothing. What this adds is the flags that do:

Flag aria2 default Why
--retry-wait=3 0 retrying with no delay burns all five tries in under a second
--timeout=30 / --connect-timeout=15 60 a dead connection otherwise holds a slot for a full minute
--lowest-speed-limit=50K 0 (off) without it a wedged connection hangs forever
--disk-cache=64M 16M sixteen writers at out-of-order offsets thrash a small cache
--file-allocation=falloc none instant on btrfs/ext4/xfs; avoids fragmenting multi-GB files
--log=logs/aria2.log none the console level hides everything, leaving no forensics

--continue is deliberately never passed — see the .part note above.

aria2c only handles http/https/ftp/ftps, so HLS and DASH fall through to yt-dlp's native downloader; --concurrent-fragments is the knob there. A missing aria2c downgrades with a warning rather than failing.

Because yt-dlp calls an external downloader exactly once and does not retry it, the retry loop lives here instead.

Options worth knowing

--formats txt,srt,vtt,json · --language en · --task translate · --playlist (a URL carrying &list= stays one video unless you ask) · --audio-profile whisper (16 kHz mono FLAC instead of source quality) · --force / --force-stage · --cookies-from-browser firefox · --js-runtime · -v

Requirements

ffmpeg, ffprobe, and optionally aria2c. Everything else comes from uv sync, including the deno binary YouTube needs for its signature challenges — yt-dlp defaults to deno only, and installing it as a dependency means it works under cron without any PATH setup.

Development

uv sync
pre-commit install --hook-type pre-commit --hook-type pre-push --hook-type commit-msg
uv run pytest

Commits run ruff, pyright and the full suite at 100% coverage; pushes run pylint and mypy. No unit test touches the GPU, the network or a real model, which is what keeps the commit gate fast. Live checks sit behind CCN_LIVE=1.

Conventional commits are enforced.

S
Description
No description provided
Readme
280 KiB
Languages
Python 99.7%
Shell 0.3%