Commit Graph
17 Commits
Author SHA1 Message Date
JMR-devandClaude Opus 5 c844781d1e fix(backends): explain why an explicitly requested device cannot be used
--device nvidia on a machine without one reported "no usable backend remains",
which says nothing about the cause and repeated the same message once per URL.
Asking for a device that cannot run is a configuration mistake, so it now fails
once, fatally, naming the reason: "nvidia/cuda: no nvidia hardware detected".

Found by end-to-end testing; the unit tests only covered the runtime-failure
path, where TranscribeError remains correct.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 15:44:43 -05:00
JMR-devandClaude Opus 5 fab20e1d32 docs: add a README covering resume, retention and hardware routing
Explains the parts that are non-obvious from the code: why artifacts rather
than a stored stage marker are the source of truth, why a .part without its
aria2 control file is discarded, which aria2c flags actually change behaviour
versus which are already yt-dlp defaults, and why NVIDIA and AMD ship as
interfaces rather than implementations.

README-intel.md keeps the hardware setup; this covers the pipeline.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 15:29:41 -05:00
JMR-devandClaude Opus 5 cdf18d0a51 feat(cli): add the command line, batch runner and doctor
doctor turns README-intel.md's troubleshooting table into code, kept two-phase
so it can distinguish a driver problem from a plugin problem. On this machine it
surfaces the two things nothing else would: the JS runtime YouTube needs, and
that /dev/dri/renderD128 is world-writable while the user is not in the render
group -- so a udev change would silently drop inference to CPU.

The retention prompt fires once per run and refuses on a non-TTY rather than
auto-accepting: a piped or cron invocation would otherwise delete gigabytes with
nobody having seen the warning. Its wording is accurate about resume still
working, because a warning users can disprove is one they stop reading.

A per-job file handler is attached and removed around each job; a 200-URL batch
would otherwise leak 200 open handlers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 15:17:45 -05:00
JMR-devandClaude Opus 5 605275e37e feat(jobs): add run configuration and the per-job runner
Retention deletions are sequenced so the guarantee actually holds: the video
goes only after the FLAC is extracted, verified and committed, and the FLAC only
after every output is written and committed. Each deletion records itself both
in state.json and in events.jsonl, so a later run can tell "removed as
requested" from "gone and should not be".

Two bugs the tests caught while writing this:

- params_changed compared the requested model id against the resolved one, so
  every re-run looked like a parameter change and re-transcribed. An unpinned
  model now matches whatever ran; only an explicit --model can disagree.
- chain=DEFAULT_CHAIN as a default argument binds at import time, which made
  the backend chain impossible to override. It is resolved at call time now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 15:04:43 -05:00
JMR-devandClaude Opus 5 1ae8c2fc96 feat(stages): add audio extraction, transcription and output writing
Audio is extracted into the job's tmp/ and moved into place only after ffprobe
confirms a readable duration, so an interrupted run never leaves a
plausible-looking stub that later verifies as complete.

transcript.json is always written even when not requested: it doubles as the
segment cache, which is what lets a deleted subtitle file be re-rendered
instead of re-transcribing a two-hour recording.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:51:04 -05:00
JMR-devandClaude Opus 5 cc1aa9a269 feat(stages): add the download stage with retry and error classification
yt-dlp messages are matched against a regex table kept as data, so the
classification is unit-testable without the network and can be amended when
YouTube changes its wording. Only network failures and aria2 process failures
retry; a private video will still be private on the third attempt, and disk
exhaustion is fatal rather than per-job.

The retry loop is ours because it has to be: ExternalFD.real_download calls the
downloader exactly once and turns a nonzero exit into report_error(), so
yt-dlp's own `retries` never covers aria2c. Since --lowest-speed-limit
deliberately makes aria2c exit nonzero on a stall, without this loop that stall
would end the job instead of resuming it -- the .aria2 control file survives, so
each attempt picks up where the last stopped.

A missing aria2c downgrades to the native downloader with a warning rather than
failing the run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:48:09 -05:00
JMR-devandClaude Opus 5 3d511f8fea feat(media): build yt-dlp and aria2c options
Only flags that actually change aria2's behaviour are sent. yt-dlp already
passes -x16 -j16 -s16 --min-split-size 1M, and -x16 is aria2's per-server cap,
so re-sending connection counts would be a no-op. What is added: a retry-wait
(aria2 defaults to 0 and burns all five tries in under a second), shorter
timeouts, a 64M disk cache for sixteen out-of-order writers, falloc on
filesystems where it is instant (this one is btrfs), a file log because
--console-log-level=warn leaves no forensics, and --lowest-speed-limit, which
is off by default and is what stops a wedged connection hanging forever.

--continue is deliberately never passed: aria2 writes segments out of order,
so resuming from a .part's length would read a sparse file as a valid prefix
and silently corrupt the video.

external_downloader is mapped to exactly the four protocols Aria2cFD supports
rather than "default", so HLS and DASH visibly fall through to the native
downloader and concurrent_fragment_downloads is what matters there.

noplaylist defaults to true so a video URL carrying &list= stays one video,
which is the opposite of yt-dlp's own default.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:43:07 -05:00
JMR-devandClaude Opus 5 2ca46074d3 feat(backends): add GPU detection, runtime fallback and the OpenVINO backend
Detection is two-phase by construction: the registry only calls probe_toolchain
once probe_hardware confirms the vendor, which is what structurally keeps torch
from being imported on an Intel-only box. A test asserts exactly that.

Hardware probing reads each render node's bound driver rather than loaded kernel
modules: /sys/module/xe exists here with zero bound devices while i915 owns the
card, so a module-presence check false-positives.

Fallback is runtime, not detection-time -- construction and the first generate()
sit in the same try, because the render-node permission failure and the OpenCL
JIT failure both surface there rather than at device enumeration. A failure
demotes the backend process-wide so a 50-job batch does not retry it 50 times,
and an explicit --device never falls back silently.

NVIDIA and AMD are interface-only: detection is real and the error names the
module to implement and the model format required. The cached model is OpenVINO
IR and cannot load on CUDA or ROCm, and CTranslate2 has no ROCm support, so
those are two separate paths rather than one parameterized one.

Model resolution is offline-first, and CACHE_DIR is anchored under XDG rather
than the working directory, which the benchmark scripts depend on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:33:21 -05:00
JMR-devandClaude Opus 5 2cfd7ad087 feat(jobs): add job identity, state, verification and resume planning
Artifacts on disk are the source of truth. state.json deliberately has no
top-level "stage" field -- persisting one is how "marked done but the file
is gone" bugs happen -- so the resume point is computed from what verifies.

The distinction that makes --no-retain safe is deleted_by_policy vs missing.
verify_* short-circuits on a policy deletion before touching the filesystem,
because probing a deliberately absent file would raise and degrade the whole
feature into "re-download everything".

A .part without its .aria2 control file is treated as unresumable: aria2
writes segments out of order, so such a file is sparse with holes rather
than a valid prefix, and resuming from its length yields a corrupt video.

Planning walks stages backwards. A policy deletion satisfies a stage that is
not re-running, but not one that is -- so --force-stage transcribe correctly
walks back to re-download. Saved segments let a deleted subtitle file be
re-rendered without re-transcribing a long recording.

state.json is written tmp -> fsync -> replace -> fsync(dir), with a test that
a failed replace leaves the previous record intact.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:24:24 -05:00
JMR-devandClaude Opus 5 15a28495b8 chore(tooling): enable flake8-bandit rules
Turns on ruff's S rules, which matter for code that shells out to ffmpeg,
aria2c and yt-dlp. S607 is ignored project-wide: binaries are looked up on
PATH deliberately and preflight-checked with shutil.which, so hardcoding
absolute paths would be less portable rather than safer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:17:41 -05:00
JMR-devandClaude Opus 5 5042d73770 feat(media): add ffmpeg decode, probe and FLAC extraction helpers
Lifts decode() out of scripts/smoke.py, where it was duplicated verbatim in
scripts/bench.py, and adds the probe and extraction calls the pipeline needs.

Command construction is separate from execution so argument lists can be
asserted directly instead of by monkeypatching internals. flac_command always
passes -sample_fmt s16: the FLAC encoder accepts only s16/s32 while AAC and
Opus decode to fltp, so leaving it to filter-graph negotiation can fail.

The default profile keeps the source rate and channels, since lossless
extraction is the point of choosing FLAC; downmixing to 16 kHz mono happens
at decode time instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:16:08 -05:00
JMR-devandClaude Opus 5 56a54c2b9a feat(transcript): add segment model, normalization and output writers
normalize_segments repairs what Whisper actually emits: an end timestamp of
-1 when the closing token is never predicted (usually the trailing chunk),
out-of-bounds and inverted spans, and blank text. Left alone these produce
malformed subtitles.

It also warns when chunk starts are non-monotonic, which is the cheap signal
that start_ts is window-relative rather than absolute -- that would misplace
every cue past 0:30, and it should surface as a log line rather than a user
report.

Timestamps convert to integer milliseconds before splitting into fields;
formatting the seconds field directly renders 3599.9996 as "00:59:60.000".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:12:39 -05:00
JMR-devandClaude Opus 5 157e846490 chore(tooling): keep ruff format out of Markdown, allow 10 dataclass fields
ruff format restyles Python blocks inside Markdown, which silently rewrote
the ffmpeg sample in README-intel.md. Documentation is not ours to restyle,
so exclude *.md.

Raise pylint max-attributes to 10: TranscriptResult is a data model and its
eight fields are not a design smell.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:12:38 -05:00
JMR-devandClaude Opus 5 c86749973c style(scripts): apply ruff formatting to the benchmark scripts
Formatting only: split combined imports and semicolon statements, sort
imports, wrap subprocess argument lists. No behaviour change; the
one-device-per-process structure that makes bench.py's numbers
trustworthy is untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:07:58 -05:00
JMR-devandClaude Opus 5 b2d648e532 feat(errors): add exception hierarchy with actionable hints
Splits failures into FatalError (abort the run) and JobError (record and
continue the batch), which is the distinction the batch runner needs.
DiskFull and SchemaTooNew are fatal on purpose: continuing past ENOSPC
only produces more corrupt artifacts.

Only NetworkError is retryable; retrying a private video is noise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:07:17 -05:00
JMR-devandClaude Opus 5 7a62688355 chore(tooling): set up uv, ruff, pyright, pylint, mypy and git hooks
Replace the ad-hoc stdlib venv with a uv-managed, locked environment and
add the quality gates:

- pre-commit: ruff (lint + format), pyright, pytest at 100% coverage
- pre-push:   pylint, mypy --strict
- commit-msg: conventional commits

pytest and pyright run with pass_filenames: false so they always see the
whole project; --cov-fail-under stays in the hook rather than addopts so
single-file TDD runs are not blocked by coverage.

requires-python is >=3.12 because numpy 2.5 does not support 3.11.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:06:38 -05:00
JMR-devandClaude Opus 5 35b8adfa98 Add Intel GPU setup docs and transcription benchmark scripts
Set up this workstation to run Whisper on the Intel Iris Xe iGPU via
OpenVINO GenAI, and capture the setup steps and measured baseline.

- README-intel.md: end-to-end setup for Intel GPUs on Linux — compute
  runtime install, render-node permissions, venv, pre-converted models,
  verification, and troubleshooting.
- scripts/smoke.py: transcribes a clip on GPU and CPU, reports timings.
- scripts/bench.py: 3 passes per device in an isolated process, also
  reporting CPU-time consumed to quantify offload.

Measured on Iris Xe (80 EU) + i5-1145G7 with large-v3-turbo-int8 over
121s of audio: GPU ~23s (~5.2x realtime, 1.0 cores busy) vs CPU ~39s
(~3.1x, 3.8 cores busy) — ~1.7x faster using ~6x less CPU time.

Two findings recorded in the README because both silently mislead:
Python 3.14 defaults multiprocessing to forkserver, so an unguarded
script runs a second copy of itself concurrently and inflates timings;
and short clips are dominated by fixed overhead, where GPU and CPU tie.

No pipeline code yet — environment setup only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 13:06:53 -05:00