The distribution, console script and import package are now audio-scribe /
audio_scribe (src/audio_scribe). Everything named for the old project follows:
- CcnError -> AudioScribeError, and its code "ccn_error" -> "audio_scribe_error"
(the code is never persisted, so existing job state still loads)
- CCN_LIVE -> AUDIO_SCRIBE_LIVE for the live-GPU tests
- OpenVINO kernel cache moves to <cache>/audio-scribe/ov_cache; the first run
after upgrading recompiles kernels, and the old directory is left in place
- README, build-binary.sh, hatch/coverage config and uv.lock updated to match
Breaking: the command is now `audio-scribe`; reinstall any tool install of the
old name with `uv tool uninstall ccn-transcribe && uv tool install .`.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Depending on yt-dlp[deno] puts a deno binary in the environment's bin
directory, and the code assumed yt-dlp would then find it. That only holds
while that directory is on PATH, which is true under `uv run` in the checkout
and false once the package is installed as a tool: only the console script is
exposed. Installed that way, doctor reported "only nvm node ... invisible to
cron" while a 95 MB deno sat unused in the tool's own bin, and YouTube
extraction was one nvm-less environment away from failing outright -- the exact
failure the deno dependency was added to prevent.
Runtimes are now resolved against sys.prefix before PATH and passed to yt-dlp
as an absolute path, which its js_runtimes config documents. That also removes
the special case that made node the only runtime resolved to a real path.
The doctor tests hid this: they faked absence by patching shutil.which alone,
so the real deno was found regardless and one of them passed for the wrong
reason. They now hide it from both lookups.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Resolving a URL to a job id cost a metadata request on every invocation,
including re-runs with nothing to do. index.json records what the last
expansion produced and a single-video URL is now answered from it: measured on
a finished job, 2.25s and three requests becomes 0.75s and none, and the resume
succeeds with the network blocked entirely, where it previously failed.
A playlist is deliberately still re-read. Answering one from the index would
silently ignore videos added to it since the last run, and that request is the
only way to notice them. Entries are keyed by whether --playlist was in effect,
because the same URL names a different set with and without it.
When a URL that has been expanded before can no longer be read, the recorded
jobs are used and the reason is logged rather than failing the whole URL. This
keys on the source being unreadable rather than on a particular error, because
an unreachable host surfaces as "no metadata" -- yt-dlp reports the failure and
returns nothing rather than raising. Only a previously expanded URL can reach
this, so it cannot invent work, and any job still needing a download fails on
its own merits. That NotFoundError's hint no longer claims a bad URL is the
only explanation.
Target replaces the (url, entry) tuple so a job id survives without metadata,
and store grows the atomic JSON write that save and the index now share.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
store.load already logged "rebuilding from artifacts" on a damaged state.json,
but nothing rebuilt anything: run_job fell back to JobState.new, so every
artifact path was None and verification reported MISSING for files sitting on
disk. Measured on a finished job with video, audio and all four transcripts
present, a one-byte corruption of state.json planned
"download -> audio -> transcribe -> outputs". It now plans nothing.
recover.rebuild scans the job directory for the merged video, the FLAC and the
written outputs, then recovers the two things the filesystem cannot answer:
events.jsonl says *why* an artifact is absent, so a policy deletion is not
mistaken for data loss and re-fetched, and transcript.json says how the
transcript was produced, so a complete job is not re-transcribed. A deletion
event for a file that is present again is treated as history, not truth.
transcript.json gains task and num_beams. It already claimed to record how the
transcript was made while omitting two parameters that change the result, and
recovery needs them to decide whether a re-run is warranted. Transcripts written
before this default to transcribe/1, which can only cost a re-transcription
rather than mislabel one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three defects found by truncating a real download and killing one mid-flight.
Matroska and WebM write the duration into the header, so ffprobe reports the
full 212.8s of a video.mkv cut off after a kilobyte -- exit 0, plausible
answer. verify_video therefore returned OK, the corrupt file was never
re-downloaded, ffmpeg extracted the 0.02s of audio it could find, and the run
reported "1 ok" with an empty transcript. The size recorded at download time
is the only evidence the bytes are still there, so it is now checked whenever
it is known rather than only as a fallback when the duration is unreadable.
The audio stage now takes the expected duration and rejects an extraction that
does not match it. ffmpeg exits 0 on a truncated container, so without this a
damaged source yields a confident transcript of near-silence, which is a worse
outcome than a failed job. It compares against the video's own probed duration
rather than state.duration_s, which can come from playlist metadata.
aria2 saves its control file every 60s by default. Since that file is what a
resume reads, a kill -9 inside the first minute preserved a control file
recording zero completed pieces: measured 0/13 on the sample, so the "resume"
re-downloaded the lot while reporting a partial. At --auto-save-interval=20 the
same kill preserves 2/13 pieces and the resumed download is byte-identical to a
clean one.
The duration tolerance moves to media/ffmpeg.py, which both callers already
import, instead of being restated per call site.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two cases the resume plan already routed to the download stage, but which the
stage could not actually fix.
yt-dlp's overwrites default is off (existing_video_file passes
default_overwrite=False), so an existing target is reported as "already
downloaded" and skipped. A truncated video.mkv was therefore planned for
re-download, never re-downloaded, and re-probed as the same corrupt file --
a job that reports success on media it could not read.
A .part without its .aria2 control file is a sparse file, not a valid prefix,
because aria2 writes its segments out of order. continuedl would resume from
its length and produce a video that silently stops being the source partway
through. Re-downloading is the cheaper mistake, so orphans go unconditionally;
a .part with its control file is left for aria2 to resume, and the sweep runs
once before the retry loop rather than between attempts.
find_video now matches only the merged output. Per-format streams from a
failed merge sort ahead of video.mkv and hold one stream each, so it could
return a video-only file to the audio stage.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The flag was parsed into RunConfig and never read, so it silently had no
effect. Finished transcripts are now also copied into that folder as
<title>-<job_id>.<ext>, which is what makes a batch readable without
digging through job directories.
It copies rather than moves, and only the formats actually requested: the
job directory stays canonical because resume depends on it, and
transcript.json is written as the segment cache whether or not it was asked
for.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The pre-push pylint gate flagged it; the score was already 10.00 but the hook
exits nonzero on any message, which is the behaviour we want.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
--device nvidia on a machine without one reported "no usable backend remains",
which says nothing about the cause and repeated the same message once per URL.
Asking for a device that cannot run is a configuration mistake, so it now fails
once, fatally, naming the reason: "nvidia/cuda: no nvidia hardware detected".
Found by end-to-end testing; the unit tests only covered the runtime-failure
path, where TranscribeError remains correct.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
doctor turns README-intel.md's troubleshooting table into code, kept two-phase
so it can distinguish a driver problem from a plugin problem. On this machine it
surfaces the two things nothing else would: the JS runtime YouTube needs, and
that /dev/dri/renderD128 is world-writable while the user is not in the render
group -- so a udev change would silently drop inference to CPU.
The retention prompt fires once per run and refuses on a non-TTY rather than
auto-accepting: a piped or cron invocation would otherwise delete gigabytes with
nobody having seen the warning. Its wording is accurate about resume still
working, because a warning users can disprove is one they stop reading.
A per-job file handler is attached and removed around each job; a 200-URL batch
would otherwise leak 200 open handlers.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Retention deletions are sequenced so the guarantee actually holds: the video
goes only after the FLAC is extracted, verified and committed, and the FLAC only
after every output is written and committed. Each deletion records itself both
in state.json and in events.jsonl, so a later run can tell "removed as
requested" from "gone and should not be".
Two bugs the tests caught while writing this:
- params_changed compared the requested model id against the resolved one, so
every re-run looked like a parameter change and re-transcribed. An unpinned
model now matches whatever ran; only an explicit --model can disagree.
- chain=DEFAULT_CHAIN as a default argument binds at import time, which made
the backend chain impossible to override. It is resolved at call time now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Audio is extracted into the job's tmp/ and moved into place only after ffprobe
confirms a readable duration, so an interrupted run never leaves a
plausible-looking stub that later verifies as complete.
transcript.json is always written even when not requested: it doubles as the
segment cache, which is what lets a deleted subtitle file be re-rendered
instead of re-transcribing a two-hour recording.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
yt-dlp messages are matched against a regex table kept as data, so the
classification is unit-testable without the network and can be amended when
YouTube changes its wording. Only network failures and aria2 process failures
retry; a private video will still be private on the third attempt, and disk
exhaustion is fatal rather than per-job.
The retry loop is ours because it has to be: ExternalFD.real_download calls the
downloader exactly once and turns a nonzero exit into report_error(), so
yt-dlp's own `retries` never covers aria2c. Since --lowest-speed-limit
deliberately makes aria2c exit nonzero on a stall, without this loop that stall
would end the job instead of resuming it -- the .aria2 control file survives, so
each attempt picks up where the last stopped.
A missing aria2c downgrades to the native downloader with a warning rather than
failing the run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Only flags that actually change aria2's behaviour are sent. yt-dlp already
passes -x16 -j16 -s16 --min-split-size 1M, and -x16 is aria2's per-server cap,
so re-sending connection counts would be a no-op. What is added: a retry-wait
(aria2 defaults to 0 and burns all five tries in under a second), shorter
timeouts, a 64M disk cache for sixteen out-of-order writers, falloc on
filesystems where it is instant (this one is btrfs), a file log because
--console-log-level=warn leaves no forensics, and --lowest-speed-limit, which
is off by default and is what stops a wedged connection hanging forever.
--continue is deliberately never passed: aria2 writes segments out of order,
so resuming from a .part's length would read a sparse file as a valid prefix
and silently corrupt the video.
external_downloader is mapped to exactly the four protocols Aria2cFD supports
rather than "default", so HLS and DASH visibly fall through to the native
downloader and concurrent_fragment_downloads is what matters there.
noplaylist defaults to true so a video URL carrying &list= stays one video,
which is the opposite of yt-dlp's own default.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Detection is two-phase by construction: the registry only calls probe_toolchain
once probe_hardware confirms the vendor, which is what structurally keeps torch
from being imported on an Intel-only box. A test asserts exactly that.
Hardware probing reads each render node's bound driver rather than loaded kernel
modules: /sys/module/xe exists here with zero bound devices while i915 owns the
card, so a module-presence check false-positives.
Fallback is runtime, not detection-time -- construction and the first generate()
sit in the same try, because the render-node permission failure and the OpenCL
JIT failure both surface there rather than at device enumeration. A failure
demotes the backend process-wide so a 50-job batch does not retry it 50 times,
and an explicit --device never falls back silently.
NVIDIA and AMD are interface-only: detection is real and the error names the
module to implement and the model format required. The cached model is OpenVINO
IR and cannot load on CUDA or ROCm, and CTranslate2 has no ROCm support, so
those are two separate paths rather than one parameterized one.
Model resolution is offline-first, and CACHE_DIR is anchored under XDG rather
than the working directory, which the benchmark scripts depend on.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Artifacts on disk are the source of truth. state.json deliberately has no
top-level "stage" field -- persisting one is how "marked done but the file
is gone" bugs happen -- so the resume point is computed from what verifies.
The distinction that makes --no-retain safe is deleted_by_policy vs missing.
verify_* short-circuits on a policy deletion before touching the filesystem,
because probing a deliberately absent file would raise and degrade the whole
feature into "re-download everything".
A .part without its .aria2 control file is treated as unresumable: aria2
writes segments out of order, so such a file is sparse with holes rather
than a valid prefix, and resuming from its length yields a corrupt video.
Planning walks stages backwards. A policy deletion satisfies a stage that is
not re-running, but not one that is -- so --force-stage transcribe correctly
walks back to re-download. Saved segments let a deleted subtitle file be
re-rendered without re-transcribing a long recording.
state.json is written tmp -> fsync -> replace -> fsync(dir), with a test that
a failed replace leaves the previous record intact.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Turns on ruff's S rules, which matter for code that shells out to ffmpeg,
aria2c and yt-dlp. S607 is ignored project-wide: binaries are looked up on
PATH deliberately and preflight-checked with shutil.which, so hardcoding
absolute paths would be less portable rather than safer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Lifts decode() out of scripts/smoke.py, where it was duplicated verbatim in
scripts/bench.py, and adds the probe and extraction calls the pipeline needs.
Command construction is separate from execution so argument lists can be
asserted directly instead of by monkeypatching internals. flac_command always
passes -sample_fmt s16: the FLAC encoder accepts only s16/s32 while AAC and
Opus decode to fltp, so leaving it to filter-graph negotiation can fail.
The default profile keeps the source rate and channels, since lossless
extraction is the point of choosing FLAC; downmixing to 16 kHz mono happens
at decode time instead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
normalize_segments repairs what Whisper actually emits: an end timestamp of
-1 when the closing token is never predicted (usually the trailing chunk),
out-of-bounds and inverted spans, and blank text. Left alone these produce
malformed subtitles.
It also warns when chunk starts are non-monotonic, which is the cheap signal
that start_ts is window-relative rather than absolute -- that would misplace
every cue past 0:30, and it should surface as a log line rather than a user
report.
Timestamps convert to integer milliseconds before splitting into fields;
formatting the seconds field directly renders 3599.9996 as "00:59:60.000".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Splits failures into FatalError (abort the run) and JobError (record and
continue the batch), which is the distinction the batch runner needs.
DiskFull and SchemaTooNew are fatal on purpose: continuing past ENOSPC
only produces more corrupt artifacts.
Only NetworkError is retryable; retrying a private video is noise.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Replace the ad-hoc stdlib venv with a uv-managed, locked environment and
add the quality gates:
- pre-commit: ruff (lint + format), pyright, pytest at 100% coverage
- pre-push: pylint, mypy --strict
- commit-msg: conventional commits
pytest and pyright run with pass_filenames: false so they always see the
whole project; --cov-fail-under stays in the hook rather than addopts so
single-file TDD runs are not blocked by coverage.
requires-python is >=3.12 because numpy 2.5 does not support 3.11.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>