203 lines
8.9 KiB
Markdown
203 lines
8.9 KiB
Markdown
# audio-scribe
|
|
|
|
Download a video, extract its audio losslessly as FLAC, and transcribe it with
|
|
Whisper on whatever GPU the machine actually has.
|
|
|
|
Resumable from any completed stage, with errors that tell you what to do next.
|
|
|
|
```bash
|
|
uv sync
|
|
uv run audio-scribe doctor # check the environment first
|
|
uv run audio-scribe run 'https://youtu.be/...'
|
|
uv run audio-scribe run -i urls.txt --formats txt,srt
|
|
```
|
|
|
|
Hardware setup for the Intel path is documented separately in
|
|
[README-intel.md](README-intel.md).
|
|
|
|
## What it does
|
|
|
|
```
|
|
yt-dlp (aria2c) -> ffmpeg -> Whisper -> txt / srt / vtt / json
|
|
video.mkv audio.flac OpenVINO GenAI
|
|
```
|
|
|
|
Each video becomes one job directory under `--workdir` (default `transcripts/`):
|
|
|
|
```
|
|
transcripts/index.json <- which jobs each URL expanded to
|
|
transcripts/jobs/youtube-<id>/
|
|
state.json media/video.mkv out/transcript.{txt,srt,vtt,json}
|
|
events.jsonl media/audio.flac logs/{job,aria2}.log
|
|
```
|
|
|
|
## Resume
|
|
|
|
**Artifacts on disk are the source of truth.** `state.json` records only what the
|
|
filesystem cannot answer — chiefly *why* a file is absent. There is deliberately
|
|
no stored "current stage": persisting one is how "marked done but the file is
|
|
gone" bugs happen. Every run verifies what is actually present (`ffprobe`
|
|
duration > 0, non-empty, audio length matching the video) and enters at the first
|
|
stage whose output does not check out.
|
|
|
|
| Transcript | FLAC | Video | Runs |
|
|
|---|---|---|---|
|
|
| missing | ok | anything | transcribe, outputs |
|
|
| missing | missing | ok | extract, transcribe, outputs |
|
|
| missing | missing | missing | everything |
|
|
| ok | — | — | nothing |
|
|
|
|
Killing a run mid-download resumes rather than restarting, because yt-dlp keeps
|
|
its `.part` and aria2 keeps its control file. How much survives depends on the
|
|
checkpoint interval, not on the size of the `.part`: with `--file-allocation=falloc`
|
|
the file is full-size from the first second, and only the pieces recorded in the
|
|
control file are skipped. A `.part` **without** that control file is discarded and
|
|
restarted, because aria2 writes segments out of order, so such a file is sparse
|
|
with holes rather than a valid prefix -- resuming from its length would produce a
|
|
video that stops being the source partway through.
|
|
|
|
A video that exists but fails verification is unlinked before re-downloading.
|
|
yt-dlp's overwrite default is off, so it would otherwise report the corrupt file
|
|
as "already downloaded" and skip it forever. Verification checks the recorded byte
|
|
count as well as the duration: Matroska and WebM write the duration into the
|
|
header, so a file truncated to a kilobyte still probes as its full length.
|
|
|
|
Deleting one subtitle file re-renders it from the segments saved in
|
|
`transcript.json` instead of re-transcribing the whole recording.
|
|
|
|
`state.json` is a convenience, not a dependency. If it is unreadable the record
|
|
is rebuilt: the job directory is scanned for the video, the FLAC and the written
|
|
outputs, `events.jsonl` supplies *why* an artifact is absent so a deletion you
|
|
asked for is not mistaken for data loss, and `transcript.json` supplies how the
|
|
transcript was produced so a finished job is not transcribed again. Delete it on
|
|
a completed job and the next run still reports "already complete".
|
|
|
|
`index.json` maps each URL to the jobs it expanded to. Re-running a single video
|
|
is therefore answered from disk with no metadata request at all — it works with
|
|
no network — and a URL that has been seen before still resumes when the source
|
|
has become unreadable, saying so. A playlist is always re-read, since that
|
|
request is the only way to notice videos added to it since the last run.
|
|
|
|
`--dry-run` prints the plan for every job without touching anything.
|
|
|
|
## Keeping or discarding media
|
|
|
|
Video and audio are kept by default.
|
|
|
|
| Flag | Effect |
|
|
|---|---|
|
|
| `--no-retain-video` | delete the video once the job is complete |
|
|
| `--no-retain-audio` | delete the FLAC once the job is complete |
|
|
| `--no-retain` | both |
|
|
|
|
These are cleanup flags: deletion is the very last step of a job, after every
|
|
stage has finished and the outputs are published. A job that fails or is
|
|
interrupted keeps its media, so resume still works. What you actually give up is
|
|
re-run flexibility: re-transcribing later with a different `--language`, `--model`
|
|
or `--task` will need another download.
|
|
|
|
These flags prompt once per run. On a non-TTY they refuse outright rather than
|
|
proceeding — a piped or cron invocation should never delete gigabytes with nobody
|
|
seeing the warning. Pass `--yes` to confirm non-interactively.
|
|
|
|
## Hardware routing
|
|
|
|
Detection runs in two phases: a cheap hardware probe that imports nothing, then a
|
|
toolchain probe that may import. The toolchain probe is only reached once the
|
|
hardware is confirmed, which is what keeps torch off a machine with no NVIDIA or
|
|
AMD GPU.
|
|
|
|
Order is **NVIDIA → AMD → Intel GPU → CPU**, and `--device` overrides it.
|
|
|
|
Fallback happens at runtime, not at detection: constructing the backend and its
|
|
first `generate()` are both guarded, because a `/dev/dri` permission failure and
|
|
an OpenCL JIT failure surface there rather than at device enumeration. A backend
|
|
that fails is demoted for the rest of the process, so a long batch does not retry
|
|
it once per job. An explicit `--device` never falls back silently — getting the
|
|
CPU when you asked for the GPU would make any measurement a lie.
|
|
|
|
**NVIDIA and AMD are interface-only in this build.** Detection and routing are
|
|
real and the error says what to implement, but there is no inference code: the
|
|
cached model is OpenVINO IR, which cannot load on CUDA or ROCm, and neither path
|
|
could be tested here. CTranslate2 has no ROCm support, so they are genuinely two
|
|
implementations rather than one.
|
|
|
|
## Download speed
|
|
|
|
yt-dlp already passes aria2c `-x16 -j16 -s16 --min-split-size 1M`, and `-x16` is
|
|
aria2's per-server cap, so re-sending connection counts changes nothing. What
|
|
this adds is the flags that do:
|
|
|
|
| Flag | aria2 default | Why |
|
|
|---|---|---|
|
|
| `--retry-wait=3` | 0 | retrying with no delay burns all five tries in under a second |
|
|
| `--timeout=30` / `--connect-timeout=15` | 60 | a dead connection otherwise holds a slot for a full minute |
|
|
| `--lowest-speed-limit=50K` | 0 (off) | without it a wedged connection hangs forever |
|
|
| `--disk-cache=64M` | 16M | sixteen writers at out-of-order offsets thrash a small cache |
|
|
| `--file-allocation=falloc` | none | instant on btrfs/ext4/xfs; avoids fragmenting multi-GB files |
|
|
| `--auto-save-interval=20` | 60 | the control file is what a resume reads; a kill inside the first minute otherwise preserves zero completed pieces |
|
|
| `--log=logs/aria2.log` | none | the console level hides everything, leaving no forensics |
|
|
|
|
`--continue` is deliberately never passed — see the `.part` note above.
|
|
|
|
aria2c only handles `http/https/ftp/ftps`, so HLS and DASH fall through to
|
|
yt-dlp's native downloader; `--concurrent-fragments` is the knob there. A missing
|
|
aria2c downgrades with a warning rather than failing.
|
|
|
|
Because yt-dlp calls an external downloader exactly once and does not retry it,
|
|
the retry loop lives here instead.
|
|
|
|
## Options worth knowing
|
|
|
|
`--formats txt,srt,vtt,json` · `--language en` · `--task translate` ·
|
|
`--playlist` (a URL carrying `&list=` stays one video unless you ask) ·
|
|
`--audio-profile whisper` (16 kHz mono FLAC instead of source quality) ·
|
|
`--force` / `--force-stage` · `--cookies-from-browser firefox` ·
|
|
`--js-runtime` · `-v`
|
|
|
|
## Requirements
|
|
|
|
`ffmpeg`, `ffprobe`, and optionally `aria2c`. Everything else comes from
|
|
`uv sync`, including the deno binary YouTube needs for its signature challenges —
|
|
yt-dlp defaults to deno only, and installing it as a dependency means it works
|
|
under cron without any PATH setup.
|
|
|
|
## Installing
|
|
|
|
Either route puts `audio-scribe` on your PATH; `audio-scribe doctor` tells you
|
|
whether the machine can actually run it.
|
|
|
|
**As a tool** (needs uv; tracks nothing but what it installed):
|
|
|
|
```bash
|
|
uv tool install . # then: audio-scribe doctor
|
|
uv tool install . --reinstall # pick up later commits
|
|
uv tool install . --editable # or track the checkout instead
|
|
```
|
|
|
|
**As a self-contained directory** (needs neither Python nor uv at runtime):
|
|
|
|
```bash
|
|
scripts/build-binary.sh # then: dist/audio-scribe/audio-scribe doctor
|
|
```
|
|
|
|
~450 MB, most of it OpenVINO and its 47 runtime-loaded plugins. It is `onedir`
|
|
rather than `onefile` on purpose: `onefile` extracts the whole bundle to `/tmp`
|
|
on every launch. Move or copy the whole `audio-scribe/` directory, not just
|
|
the executable inside it. `ffmpeg`, `ffprobe` and `aria2c` are still expected on
|
|
the system either way — they are not bundled.
|
|
|
|
## Development
|
|
|
|
```bash
|
|
uv sync
|
|
pre-commit install --hook-type pre-commit --hook-type pre-push --hook-type commit-msg
|
|
uv run pytest
|
|
```
|
|
|
|
Commits run ruff, pyright and the full suite at 100% coverage; pushes run pylint
|
|
and mypy. No unit test touches the GPU, the network or a real model, which is
|
|
what keeps the commit gate fast. Live checks sit behind `AUDIO_SCRIBE_LIVE=1`.
|
|
|
|
Conventional commits are enforced.
|