Set up this workstation to run Whisper on the Intel Iris Xe iGPU via OpenVINO GenAI, and capture the setup steps and measured baseline. - README-intel.md: end-to-end setup for Intel GPUs on Linux — compute runtime install, render-node permissions, venv, pre-converted models, verification, and troubleshooting. - scripts/smoke.py: transcribes a clip on GPU and CPU, reports timings. - scripts/bench.py: 3 passes per device in an isolated process, also reporting CPU-time consumed to quantify offload. Measured on Iris Xe (80 EU) + i5-1145G7 with large-v3-turbo-int8 over 121s of audio: GPU ~23s (~5.2x realtime, 1.0 cores busy) vs CPU ~39s (~3.1x, 3.8 cores busy) — ~1.7x faster using ~6x less CPU time. Two findings recorded in the README because both silently mislead: Python 3.14 defaults multiprocessing to forkserver, so an unguarded script runs a second copy of itself concurrently and inflates timings; and short clips are dominated by fixed overhead, where GPU and CPU tie. No pipeline code yet — environment setup only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
36 lines
1.4 KiB
Python
36 lines
1.4 KiB
Python
"""One device per process: wall time + CPU time consumed (offload evidence)."""
|
|
import subprocess, sys, time, resource
|
|
import numpy as np
|
|
from huggingface_hub import snapshot_download
|
|
import openvino_genai as ov_genai
|
|
|
|
SR = 16000
|
|
|
|
|
|
def main():
|
|
dev = sys.argv[1]
|
|
audio = sys.argv[2] if len(sys.argv) > 2 else "sample.wav"
|
|
raw = subprocess.run(["ffmpeg", "-nostdin", "-loglevel", "error", "-i", audio,
|
|
"-f", "f32le", "-ac", "1", "-ar", str(SR), "-"],
|
|
capture_output=True, check=True).stdout
|
|
speech = np.frombuffer(raw, dtype=np.float32)
|
|
secs = len(speech) / SR
|
|
|
|
pipe = ov_genai.WhisperPipeline(
|
|
snapshot_download("OpenVINO/whisper-large-v3-turbo-int8-ov"),
|
|
dev, CACHE_DIR=f".ov_cache_{dev.lower()}")
|
|
|
|
for i in range(3):
|
|
r0 = resource.getrusage(resource.RUSAGE_SELF)
|
|
t = time.perf_counter()
|
|
pipe.generate(speech, task="transcribe", return_timestamps=True)
|
|
wall = time.perf_counter() - t
|
|
r1 = resource.getrusage(resource.RUSAGE_SELF)
|
|
cpu = (r1.ru_utime - r0.ru_utime) + (r1.ru_stime - r0.ru_stime)
|
|
print(f"{dev} pass{i+1}: wall {wall:6.1f}s RTF {secs/wall:.2f}x "
|
|
f"cpu-time {cpu:6.1f}s ({cpu/wall:4.1f} cores busy)", flush=True)
|
|
|
|
|
|
if __name__ == "__main__": # forkserver re-imports __main__; guard is required
|
|
main()
|