Add Intel GPU setup docs and transcription benchmark scripts

Set up this workstation to run Whisper on the Intel Iris Xe iGPU via
OpenVINO GenAI, and capture the setup steps and measured baseline.

- README-intel.md: end-to-end setup for Intel GPUs on Linux — compute
  runtime install, render-node permissions, venv, pre-converted models,
  verification, and troubleshooting.
- scripts/smoke.py: transcribes a clip on GPU and CPU, reports timings.
- scripts/bench.py: 3 passes per device in an isolated process, also
  reporting CPU-time consumed to quantify offload.

Measured on Iris Xe (80 EU) + i5-1145G7 with large-v3-turbo-int8 over
121s of audio: GPU ~23s (~5.2x realtime, 1.0 cores busy) vs CPU ~39s
(~3.1x, 3.8 cores busy) — ~1.7x faster using ~6x less CPU time.

Two findings recorded in the README because both silently mislead:
Python 3.14 defaults multiprocessing to forkserver, so an unguarded
script runs a second copy of itself concurrently and inflates timings;
and short clips are dominated by fixed overhead, where GPU and CPU tie.

No pipeline code yet — environment setup only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-09-13 13:06:53 -05:00
co-authored by Claude Opus 5
commit 35b8adfa98
4 changed files with 344 additions and 0 deletions
+35
View File
@@ -0,0 +1,35 @@
"""One device per process: wall time + CPU time consumed (offload evidence)."""
import subprocess, sys, time, resource
import numpy as np
from huggingface_hub import snapshot_download
import openvino_genai as ov_genai
SR = 16000
def main():
dev = sys.argv[1]
audio = sys.argv[2] if len(sys.argv) > 2 else "sample.wav"
raw = subprocess.run(["ffmpeg", "-nostdin", "-loglevel", "error", "-i", audio,
"-f", "f32le", "-ac", "1", "-ar", str(SR), "-"],
capture_output=True, check=True).stdout
speech = np.frombuffer(raw, dtype=np.float32)
secs = len(speech) / SR
pipe = ov_genai.WhisperPipeline(
snapshot_download("OpenVINO/whisper-large-v3-turbo-int8-ov"),
dev, CACHE_DIR=f".ov_cache_{dev.lower()}")
for i in range(3):
r0 = resource.getrusage(resource.RUSAGE_SELF)
t = time.perf_counter()
pipe.generate(speech, task="transcribe", return_timestamps=True)
wall = time.perf_counter() - t
r1 = resource.getrusage(resource.RUSAGE_SELF)
cpu = (r1.ru_utime - r0.ru_utime) + (r1.ru_stime - r0.ru_stime)
print(f"{dev} pass{i+1}: wall {wall:6.1f}s RTF {secs/wall:.2f}x "
f"cpu-time {cpu:6.1f}s ({cpu/wall:4.1f} cores busy)", flush=True)
if __name__ == "__main__": # forkserver re-imports __main__; guard is required
main()