3 Commits
Author SHA1 Message Date
JMR-devandClaude Sonnet 5 b23995f218 refactor: rename the package from ccn-transcribe to audio-scribe
The distribution, console script and import package are now audio-scribe /
audio_scribe (src/audio_scribe). Everything named for the old project follows:

- CcnError -> AudioScribeError, and its code "ccn_error" -> "audio_scribe_error"
  (the code is never persisted, so existing job state still loads)
- CCN_LIVE -> AUDIO_SCRIBE_LIVE for the live-GPU tests
- OpenVINO kernel cache moves to <cache>/audio-scribe/ov_cache; the first run
  after upgrading recompiles kernels, and the old directory is left in place
- README, build-binary.sh, hatch/coverage config and uv.lock updated to match

Breaking: the command is now `audio-scribe`; reinstall any tool install of the
old name with `uv tool uninstall ccn-transcribe && uv tool install .`.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 15:52:42 -05:00
JMR-devandClaude Opus 5 3bff193f92 fix(verify): catch a truncated container, and checkpoint aria2 often enough to resume
Three defects found by truncating a real download and killing one mid-flight.

Matroska and WebM write the duration into the header, so ffprobe reports the
full 212.8s of a video.mkv cut off after a kilobyte -- exit 0, plausible
answer. verify_video therefore returned OK, the corrupt file was never
re-downloaded, ffmpeg extracted the 0.02s of audio it could find, and the run
reported "1 ok" with an empty transcript. The size recorded at download time
is the only evidence the bytes are still there, so it is now checked whenever
it is known rather than only as a fallback when the duration is unreadable.

The audio stage now takes the expected duration and rejects an extraction that
does not match it. ffmpeg exits 0 on a truncated container, so without this a
damaged source yields a confident transcript of near-silence, which is a worse
outcome than a failed job. It compares against the video's own probed duration
rather than state.duration_s, which can come from playlist metadata.

aria2 saves its control file every 60s by default. Since that file is what a
resume reads, a kill -9 inside the first minute preserved a control file
recording zero completed pieces: measured 0/13 on the sample, so the "resume"
re-downloaded the lot while reporting a partial. At --auto-save-interval=20 the
same kill preserves 2/13 pieces and the resumed download is byte-identical to a
clean one.

The duration tolerance moves to media/ffmpeg.py, which both callers already
import, instead of being restated per call site.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 16:17:14 -05:00
JMR-devandClaude Opus 5 5042d73770 feat(media): add ffmpeg decode, probe and FLAC extraction helpers
Lifts decode() out of scripts/smoke.py, where it was duplicated verbatim in
scripts/bench.py, and adds the probe and extraction calls the pipeline needs.

Command construction is separate from execution so argument lists can be
asserted directly instead of by monkeypatching internals. flac_command always
passes -sample_fmt s16: the FLAC encoder accepts only s16/s32 while AAC and
Opus decode to fltp, so leaving it to filter-graph negotiation can fail.

The default profile keeps the source rate and channels, since lossless
extraction is the point of choosing FLAC; downmixing to 16 kHz mono happens
at decode time instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:16:08 -05:00