docs: describe state recovery and the URL index

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-09-13 17:00:11 -05:00
co-authored by Claude Opus 5
parent 1114595de7
commit 5661b49942
+14
View File
@@ -25,6 +25,7 @@ yt-dlp (aria2c) -> ffmpeg -> Whisper -> txt / srt / vtt / json
Each video becomes one job directory under `--workdir` (default `transcripts/`):
```
transcripts/index.json <- which jobs each URL expanded to
transcripts/jobs/youtube-<id>/
state.json media/video.mkv out/transcript.{txt,srt,vtt,json}
events.jsonl media/audio.flac logs/{job,aria2}.log
@@ -64,6 +65,19 @@ header, so a file truncated to a kilobyte still probes as its full length.
Deleting one subtitle file re-renders it from the segments saved in
`transcript.json` instead of re-transcribing the whole recording.
`state.json` is a convenience, not a dependency. If it is unreadable the record
is rebuilt: the job directory is scanned for the video, the FLAC and the written
outputs, `events.jsonl` supplies *why* an artifact is absent so a deletion you
asked for is not mistaken for data loss, and `transcript.json` supplies how the
transcript was produced so a finished job is not transcribed again. Delete it on
a completed job and the next run still reports "already complete".
`index.json` maps each URL to the jobs it expanded to. Re-running a single video
is therefore answered from disk with no metadata request at all — it works with
no network — and a URL that has been seen before still resumes when the source
has become unreadable, saying so. A playlist is always re-read, since that
request is the only way to notice videos added to it since the last run.
`--dry-run` prints the plan for every job without touching anything.
## Keeping or discarding media