Commit Graph
4 Commits
Author SHA1 Message Date
JMR-devandClaude Sonnet 5 b23995f218 refactor: rename the package from ccn-transcribe to audio-scribe
The distribution, console script and import package are now audio-scribe /
audio_scribe (src/audio_scribe). Everything named for the old project follows:

- CcnError -> AudioScribeError, and its code "ccn_error" -> "audio_scribe_error"
  (the code is never persisted, so existing job state still loads)
- CCN_LIVE -> AUDIO_SCRIBE_LIVE for the live-GPU tests
- OpenVINO kernel cache moves to <cache>/audio-scribe/ov_cache; the first run
  after upgrading recompiles kernels, and the old directory is left in place
- README, build-binary.sh, hatch/coverage config and uv.lock updated to match

Breaking: the command is now `audio-scribe`; reinstall any tool install of the
old name with `uv tool uninstall ccn-transcribe && uv tool install .`.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 15:52:42 -05:00
JMR-devandClaude Opus 5 e459606e89 feat(packaging): ship as a tool install or a self-contained directory
Two ways to get `ccn-transcribe` on PATH. `uv tool install .` is the light one.
scripts/build-binary.sh produces a ~450 MB directory that needs neither Python
nor uv, which turned out to work despite OpenVINO resolving 47 plugins by
dlopen at runtime -- the frozen build reports CPU, GPU and transcribes on the
GPU to a byte-identical transcript.

Two things that build needs. PyInstaller points sys.prefix at the bundle, so
deno goes in bin/ and the runtime lookup finds it with no frozen-app special
case in the package. And a frozen executable is re-invoked to start a
multiprocessing child, which parsed the interpreter's own -B flag as a CLI
option and printed "No such option '-B'" twice per run before dying;
freeze_support() in the entry point makes that child do its job and exit.

onedir, not onefile: onefile would extract 450 MB to /tmp on every launch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:10:51 -05:00
JMR-devandClaude Opus 5 c86749973c style(scripts): apply ruff formatting to the benchmark scripts
Formatting only: split combined imports and semicolon statements, sort
imports, wrap subprocess argument lists. No behaviour change; the
one-device-per-process structure that makes bench.py's numbers
trustworthy is untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:07:58 -05:00
JMR-devandClaude Opus 5 35b8adfa98 Add Intel GPU setup docs and transcription benchmark scripts
Set up this workstation to run Whisper on the Intel Iris Xe iGPU via
OpenVINO GenAI, and capture the setup steps and measured baseline.

- README-intel.md: end-to-end setup for Intel GPUs on Linux — compute
  runtime install, render-node permissions, venv, pre-converted models,
  verification, and troubleshooting.
- scripts/smoke.py: transcribes a clip on GPU and CPU, reports timings.
- scripts/bench.py: 3 passes per device in an isolated process, also
  reporting CPU-time consumed to quantify offload.

Measured on Iris Xe (80 EU) + i5-1145G7 with large-v3-turbo-int8 over
121s of audio: GPU ~23s (~5.2x realtime, 1.0 cores busy) vs CPU ~39s
(~3.1x, 3.8 cores busy) — ~1.7x faster using ~6x less CPU time.

Two findings recorded in the README because both silently mislead:
Python 3.14 defaults multiprocessing to forkserver, so an unguarded
script runs a second copy of itself concurrently and inflates timings;
and short clips are dominated by fixed overhead, where GPU and CPU tie.

No pipeline code yet — environment setup only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 13:06:53 -05:00