Detection is two-phase by construction: the registry only calls probe_toolchain
once probe_hardware confirms the vendor, which is what structurally keeps torch
from being imported on an Intel-only box. A test asserts exactly that.
Hardware probing reads each render node's bound driver rather than loaded kernel
modules: /sys/module/xe exists here with zero bound devices while i915 owns the
card, so a module-presence check false-positives.
Fallback is runtime, not detection-time -- construction and the first generate()
sit in the same try, because the render-node permission failure and the OpenCL
JIT failure both surface there rather than at device enumeration. A failure
demotes the backend process-wide so a 50-job batch does not retry it 50 times,
and an explicit --device never falls back silently.
NVIDIA and AMD are interface-only: detection is real and the error names the
module to implement and the model format required. The cached model is OpenVINO
IR and cannot load on CUDA or ROCm, and CTranslate2 has no ROCm support, so
those are two separate paths rather than one parameterized one.
Model resolution is offline-first, and CACHE_DIR is anchored under XDG rather
than the working directory, which the benchmark scripts depend on.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>