Files
JMR-devandClaude Opus 5 4d19d62077 docs: record what a resume actually preserves, and date the benchmark
The resume section implied a .part file's size reflected progress; with falloc
it is full-size from the first second and only the checkpointed pieces are
skipped. Adds the corrupt-video and truncation rules, and the new
--auto-save-interval row.

README-intel.md's 5.2x figure does not reproduce on this box today (~4.0x from
the repo's own bench.py, GPU/CPU ratio unchanged), so the table is marked as a
point-in-time measurement rather than a target to chase.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 16:19:00 -05:00

7.6 KiB

Running Whisper on an Intel GPU (Linux)

Setup for GPU-accelerated Whisper transcription on Intel integrated or discrete graphics, using OpenVINO GenAI.

Verified on: Fedora 44, kernel 7.1, Intel Iris Xe (Tiger Lake GT2, Gen12LP, 80 EU), i5-1145G7, Python 3.14.7, OpenVINO 2026.3.1.


1. Prerequisites

Requirement Notes
Intel GPU, Gen9 or newer OpenVINO's GPU plugin targets Gen9-Gen12LP and Arc. Gen8/9/11 need the legacy1 runtime packages instead of the mainline ones.
Python 3.10+ 3.14 works; openvino-genai ships cp310-cp314 wheels.
ffmpeg Used to decode audio. Avoids a librosa/numba dependency.

Check which GPU you have and that a kernel driver is bound:

lspci -nn | grep -Ei 'vga|display'      # e.g. Intel ... [8086:9a49]
lspci -k -s 00:02.0 | grep 'Kernel driver'   # expect i915 (older) or xe (Lunar Lake+)

2. Install the Intel compute runtime

This is the piece that's usually missing. OpenVINO's GPU plugin is OpenCL-based, so without it OpenVINO sees only the CPU. The Mesa/Vulkan driver that ships with your desktop is not sufficient.

Fedora / RHEL:

sudo dnf install intel-compute-runtime oneapi-level-zero clinfo

Ubuntu / Debian:

sudo apt install intel-opencl-icd libze1 clinfo

Arch:

sudo pacman -S intel-compute-runtime level-zero-loader clinfo

Package names other than the Fedora set are the usual spellings but were not verified here — check your distro's repos if they don't resolve. On older Intel hardware (Gen8/9/11) look for -legacy1 variants.

Verify — this must list your GPU before anything else will work:

clinfo -l
# Platform #0: Intel(R) OpenCL Graphics
#  `-- Device #0: Intel(R) Iris(R) Xe Graphics

Device permissions

Your user needs access to the render node. Check it:

ls -l /dev/dri/renderD128

If the mode is 0660 (rather than 0666), add yourself to the owning group and log out and back in:

sudo usermod -aG render $USER    # some distros use 'video' instead

3. Python environment

python3 -m venv .venv && source .venv/bin/activate
pip install openvino-genai huggingface-hub numpy

Install from PyPI, not your distro's python3-openvino package — distro builds tend to lag several releases behind and will conflict with the pip install.

Verify the plugin sees the GPU:

python -c "import openvino as ov; print(ov.Core().available_devices)"
# ['CPU', 'GPU']

If clinfo -l succeeded but GPU is missing here, the problem is the OpenVINO plugin rather than the driver — the two checks are deliberately separate.

4. Get a model

Use Intel's pre-converted models; this avoids pulling torch, optimum-intel and nncf (several GB) just to convert weights:

hf download OpenVINO/whisper-large-v3-turbo-int8-ov

Other sizes exist under the same org — whisper-{tiny,base,small,medium,large-v3} and distil-whisper-large-v3, each in fp16, int8 and int4. The int8 builds are weight-only compressed; on an iGPU the win is memory bandwidth, which is the binding constraint when the GPU shares system RAM with the CPU.

5. Verify end to end

Fetch a test clip and loop it so long-form (>30 s) chunking is exercised:

curl -sLO https://raw.githubusercontent.com/ggml-org/whisper.cpp/master/samples/jfk.wav
ffmpeg -stream_loop 10 -i jfk.wav -ar 16000 -ac 1 sample.wav   # ~2 min
python scripts/smoke.py sample.wav       # transcript + GPU/CPU timings
python scripts/bench.py GPU sample.wav   # 3 passes, one device per process
python scripts/bench.py CPU sample.wav

Expected results

Measured on Iris Xe (80 EU) + i5-1145G7, large-v3-turbo-int8, 121 s of audio:

device wall realtime factor CPU-time cores busy
GPU ~23 s ~5.2x ~24 s 1.0
CPU ~39 s ~3.1x ~147 s 3.8

These are a point-in-time measurement. Re-running scripts/bench.py GPU sample.wav on the same box in September 2026 gave 29-32 s (~4.0x) rather than 23 s, with the GPU/CPU ratio unchanged. Treat the ratio as the durable result and re-measure the absolute numbers on your own machine before calling a slower run a regression.

The iGPU is ~1.7x faster and uses ~6x less CPU time. Note it still occupies one full core — the plugin busy-waits — so the offload is 3.8 cores down to 1.0, not to zero. Budget for that if transcription runs alongside other CPU work.

Don't expect Arc-class numbers from an integrated part: Xe-LP has no XMX matrix engines, so the gain comes from bandwidth and offload rather than raw matrix throughput.

Benchmark with at least a minute of audio. On a short clip the fixed per-call overhead dominates and the GPU advantage disappears entirely - on an 11 s sample the same hardware measured 7.9 s on GPU vs 7.8 s on CPU, a dead heat, versus the 1.7x gap on 121 s. Short-clip numbers will tell you the GPU is pointless when it isn't.

6. Usage notes

Enable the kernel cache. The GPU plugin JIT-compiles OpenCL kernels; CACHE_DIR works as a plain keyword argument:

pipe = ov_genai.WhisperPipeline(model_dir, "GPU", CACHE_DIR=".ov_cache_gpu")
result = pipe.generate(speech, task="transcribe", return_timestamps=True)
for c in result.chunks:
    print(c.start_ts, c.end_ts, c.text)

Model load is ~7 s cold and ~0.7 s cached. Separately, the first transcription on a fresh cache costs ~31 s vs ~23 s steady-state, because shape-specific kernels are compiled lazily at first inference rather than at load. A slow first run is expected, not a fault.

Audio must be 16 kHz mono float32. Decode with ffmpeg rather than a Python audio library:

raw = subprocess.run(["ffmpeg", "-nostdin", "-loglevel", "error", "-i", path,
                      "-f", "f32le", "-ac", "1", "-ar", "16000", "-"],
                     capture_output=True, check=True).stdout
speech = np.frombuffer(raw, dtype=np.float32)

Long-form audio needs no chunking code of your own — WhisperPipeline applies a sequential 30-second sliding window internally.

Always guard your entrypoints:

if __name__ == "__main__":
    main()

Python 3.14 changed the default multiprocessing start method on Linux from fork to forkserver, which re-imports __main__. Without the guard, a script whose dependencies touch multiprocessing runs a second copy of itself concurrently. It fails silently — the symptom is duplicated output lines and timings inflated by self-contention. This cost us a benchmark that reported the GPU 3x slower than it is.

7. Troubleshooting

Symptom Cause
clinfo -l lists no platform Compute runtime not installed, or too old for your GPU.
clinfo works, OpenVINO shows only ['CPU'] Plugin problem, not a driver problem. Check you're in the venv and not shadowed by a distro python3-openvino.
Permission denied on /dev/dri/renderD128 Not in the render (or video) group; needs a re-login.
First run very slow, later runs fine Kernel JIT populating CACHE_DIR. Expected.
Output lines appear twice; everything slow Missing if __name__ == "__main__" guard — see above.

Alternative: whisper.cpp + Vulkan

If the OpenCL stack won't cooperate, whisper.cpp's Vulkan backend needs no Intel runtime at all — the Mesa driver most desktops already have is enough:

git clone https://github.com/ggml-org/whisper.cpp && cd whisper.cpp
cmake -B build -DGGML_VULKAN=ON && cmake --build build -j$(nproc)

Generally slower than OpenVINO on Intel hardware, but a useful fallback and an independent check on your numbers.