Commit Graph
10 Commits
Author SHA1 Message Date
JMR-devandClaude Opus 5 2ca46074d3 feat(backends): add GPU detection, runtime fallback and the OpenVINO backend
Detection is two-phase by construction: the registry only calls probe_toolchain
once probe_hardware confirms the vendor, which is what structurally keeps torch
from being imported on an Intel-only box. A test asserts exactly that.

Hardware probing reads each render node's bound driver rather than loaded kernel
modules: /sys/module/xe exists here with zero bound devices while i915 owns the
card, so a module-presence check false-positives.

Fallback is runtime, not detection-time -- construction and the first generate()
sit in the same try, because the render-node permission failure and the OpenCL
JIT failure both surface there rather than at device enumeration. A failure
demotes the backend process-wide so a 50-job batch does not retry it 50 times,
and an explicit --device never falls back silently.

NVIDIA and AMD are interface-only: detection is real and the error names the
module to implement and the model format required. The cached model is OpenVINO
IR and cannot load on CUDA or ROCm, and CTranslate2 has no ROCm support, so
those are two separate paths rather than one parameterized one.

Model resolution is offline-first, and CACHE_DIR is anchored under XDG rather
than the working directory, which the benchmark scripts depend on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:33:21 -05:00
JMR-devandClaude Opus 5 2cfd7ad087 feat(jobs): add job identity, state, verification and resume planning
Artifacts on disk are the source of truth. state.json deliberately has no
top-level "stage" field -- persisting one is how "marked done but the file
is gone" bugs happen -- so the resume point is computed from what verifies.

The distinction that makes --no-retain safe is deleted_by_policy vs missing.
verify_* short-circuits on a policy deletion before touching the filesystem,
because probing a deliberately absent file would raise and degrade the whole
feature into "re-download everything".

A .part without its .aria2 control file is treated as unresumable: aria2
writes segments out of order, so such a file is sparse with holes rather
than a valid prefix, and resuming from its length yields a corrupt video.

Planning walks stages backwards. A policy deletion satisfies a stage that is
not re-running, but not one that is -- so --force-stage transcribe correctly
walks back to re-download. Saved segments let a deleted subtitle file be
re-rendered without re-transcribing a long recording.

state.json is written tmp -> fsync -> replace -> fsync(dir), with a test that
a failed replace leaves the previous record intact.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:24:24 -05:00
JMR-devandClaude Opus 5 15a28495b8 chore(tooling): enable flake8-bandit rules
Turns on ruff's S rules, which matter for code that shells out to ffmpeg,
aria2c and yt-dlp. S607 is ignored project-wide: binaries are looked up on
PATH deliberately and preflight-checked with shutil.which, so hardcoding
absolute paths would be less portable rather than safer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:17:41 -05:00
JMR-devandClaude Opus 5 5042d73770 feat(media): add ffmpeg decode, probe and FLAC extraction helpers
Lifts decode() out of scripts/smoke.py, where it was duplicated verbatim in
scripts/bench.py, and adds the probe and extraction calls the pipeline needs.

Command construction is separate from execution so argument lists can be
asserted directly instead of by monkeypatching internals. flac_command always
passes -sample_fmt s16: the FLAC encoder accepts only s16/s32 while AAC and
Opus decode to fltp, so leaving it to filter-graph negotiation can fail.

The default profile keeps the source rate and channels, since lossless
extraction is the point of choosing FLAC; downmixing to 16 kHz mono happens
at decode time instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:16:08 -05:00
JMR-devandClaude Opus 5 56a54c2b9a feat(transcript): add segment model, normalization and output writers
normalize_segments repairs what Whisper actually emits: an end timestamp of
-1 when the closing token is never predicted (usually the trailing chunk),
out-of-bounds and inverted spans, and blank text. Left alone these produce
malformed subtitles.

It also warns when chunk starts are non-monotonic, which is the cheap signal
that start_ts is window-relative rather than absolute -- that would misplace
every cue past 0:30, and it should surface as a log line rather than a user
report.

Timestamps convert to integer milliseconds before splitting into fields;
formatting the seconds field directly renders 3599.9996 as "00:59:60.000".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:12:39 -05:00
JMR-devandClaude Opus 5 157e846490 chore(tooling): keep ruff format out of Markdown, allow 10 dataclass fields
ruff format restyles Python blocks inside Markdown, which silently rewrote
the ffmpeg sample in README-intel.md. Documentation is not ours to restyle,
so exclude *.md.

Raise pylint max-attributes to 10: TranscriptResult is a data model and its
eight fields are not a design smell.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:12:38 -05:00
JMR-devandClaude Opus 5 c86749973c style(scripts): apply ruff formatting to the benchmark scripts
Formatting only: split combined imports and semicolon statements, sort
imports, wrap subprocess argument lists. No behaviour change; the
one-device-per-process structure that makes bench.py's numbers
trustworthy is untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:07:58 -05:00
JMR-devandClaude Opus 5 b2d648e532 feat(errors): add exception hierarchy with actionable hints
Splits failures into FatalError (abort the run) and JobError (record and
continue the batch), which is the distinction the batch runner needs.
DiskFull and SchemaTooNew are fatal on purpose: continuing past ENOSPC
only produces more corrupt artifacts.

Only NetworkError is retryable; retrying a private video is noise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:07:17 -05:00
JMR-devandClaude Opus 5 7a62688355 chore(tooling): set up uv, ruff, pyright, pylint, mypy and git hooks
Replace the ad-hoc stdlib venv with a uv-managed, locked environment and
add the quality gates:

- pre-commit: ruff (lint + format), pyright, pytest at 100% coverage
- pre-push:   pylint, mypy --strict
- commit-msg: conventional commits

pytest and pyright run with pass_filenames: false so they always see the
whole project; --cov-fail-under stays in the hook rather than addopts so
single-file TDD runs are not blocked by coverage.

requires-python is >=3.12 because numpy 2.5 does not support 3.11.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:06:38 -05:00
JMR-devandClaude Opus 5 35b8adfa98 Add Intel GPU setup docs and transcription benchmark scripts
Set up this workstation to run Whisper on the Intel Iris Xe iGPU via
OpenVINO GenAI, and capture the setup steps and measured baseline.

- README-intel.md: end-to-end setup for Intel GPUs on Linux — compute
  runtime install, render-node permissions, venv, pre-converted models,
  verification, and troubleshooting.
- scripts/smoke.py: transcribes a clip on GPU and CPU, reports timings.
- scripts/bench.py: 3 passes per device in an isolated process, also
  reporting CPU-time consumed to quantify offload.

Measured on Iris Xe (80 EU) + i5-1145G7 with large-v3-turbo-int8 over
121s of audio: GPU ~23s (~5.2x realtime, 1.0 cores busy) vs CPU ~39s
(~3.1x, 3.8 cores busy) — ~1.7x faster using ~6x less CPU time.

Two findings recorded in the README because both silently mislead:
Python 3.14 defaults multiprocessing to forkserver, so an unguarded
script runs a second copy of itself concurrently and inflates timings;
and short clips are dominated by fixed overhead, where GPU and CPU tie.

No pipeline code yet — environment setup only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 13:06:53 -05:00