Hardware acceleration
Frameleaf can use a graphics card or other accelerator for two kinds of work:
- Machine learning in the machine learning container: smart search, faces, text recognition and AI descriptions. That’s this page.
- Video transcoding in the server container, to reduce CPU load. See Video transcoding.
Both are optional and still experimental, so they may not work on every system. The instructions are for Docker Compose on Linux, or Windows through WSL2; other container engines may need different configuration.
CPU-only processing works, but a GPU is strongly recommended for larger libraries and heavier models. You don’t need to redo any jobs after turning acceleration on. Jobs that run afterwards use the device.
Check your hardware
Section titled “Check your hardware”Settings, then Compute & jobs, then Hardware & GPU checks each part of Frameleaf separately, because each needs its own access to the GPU. For the server container (video and Studio export) and the machine learning container (AI features) it reports three facts:
| Fact | What it means |
|---|---|
| GPU on the host | The host’s PCI list, read from inside the container, shows a GPU. Under WSL2 or Docker Desktop the host’s GPUs are hidden, so this shows Not reported unless the container sees one |
| Visible to the container | An NVIDIA device, a /dev/dri render node, /dev/kfd or /dev/dxg was passed in |
| Used by the runtime | The test transcode ran on the GPU’s video engine, or the analysis runtime (CUDA, ROCm or OpenVINO) is using the GPU |
Studio render workers and restoration workers are listed below the two containers, from what they reported. Anything a worker doesn’t report shows Not reported; it’s never guessed.
The check names the set-up problems it finds, each with the fix and, where Docker Compose can fix it, a ready-to-paste snippet with the values it found, such as the render group’s number or the graphics card’s PCI path. It detects, among others:
- the NVIDIA Container Toolkit missing, or an old driver
- the
videocapability missing for NVENC /dev/drinot passed in, or a render node owned by a group the container doesn’t have- the integrated GPU used instead of the graphics card, on a host with both
- an AMD RX 6000 or 7600 card that needs
HSA_OVERRIDE_GFX_VERSIONfor ROCm - WSL2, Unraid, Docker Desktop on a Mac, and the CPU image on a host with a GPU
A Turing card (compute capability 7.5, such as the GTX 16 series) gets a note: it has no bf16 or FlashAttention, so models that use them run slower.
Benchmark
Section titled “Benchmark”Run a short benchmark times search embeddings and a test transcode, then records the speed of each kind of work the model sliders estimate:
- Descriptions and tags: a few generated test photos are described with your description model on this server’s machine learning container.
- Restoration: the speed the restoration worker measured on its GPU when it was qualified.
- Upscale has no local runner, and transcription and Smooth motion workers report no speed; the benchmark says so rather than guessing.
Supported hardware
Section titled “Supported hardware”| Backend | Hardware | Image suffix |
|---|---|---|
| CUDA | NVIDIA GPUs with compute capability 5.2 or higher. The most reliable option | -cuda |
| ROCm | AMD GPUs | -rocm |
| OpenVINO | Intel GPUs such as Iris Xe and Arc | -openvino |
| ARM NN | Devices with Mali GPUs. Other Arm devices aren’t supported | -armnn |
| RKNN | Rockchip RK3566, RK3568, RK3576 and RK3588 | -rknn |
Some models don’t work with every backend. ARM NN doesn’t speed up searching, because of model compatibility, though smart search jobs do use it.
Before you start
Section titled “Before you start”CUDA
- The official NVIDIA driver, version 545 or newer (it must support CUDA 12.3).
- On Linux (not WSL2), the NVIDIA Container Toolkit.
ROCm
- On Linux, the AMDGPU driver module. With Secure Boot, the DKMS signing key must be enrolled in the UEFI firmware.
- A GPU supported by ROCm. If yours isn’t officially supported, try
HSA_OVERRIDE_GFX_VERSION=<a supported version, such as 10.3.0>, and if that doesn’t work, alsoHSA_USE_SVM=0. - At least 35 GiB of free disk space for the ROCm image. Later updates are usually only a few hundred megabytes.
- After running, the GPU can draw more power than usual until the service has been idle for 5 minutes (set by
MACHINE_LEARNING_MODEL_TTL). The MIGraphX backend compiles models when they’re first used, so the first few results are slow.
OpenVINO
- A kernel new enough to use the device.
- Expect higher RAM use than on the CPU. Integrated GPUs are more likely to have problems than discrete ones, especially older processors or servers with little RAM.
- On WSL2, use
openvino-wsland make sure the container can see/dev/dri: rundocker compose exec -T immich-machine-learning ls -la /dev/dri. If it can’t, find the group numbers withgetent group renderandgetent group videoon the WSL host and add them undergroup_addin theopenvino-wslsection ofhwaccel.ml.yml.
ARM NN
- The vendor’s kernel driver, usually preinstalled on the vendor’s Linux image.
/dev/mali0on the host (check withls /dev).- The closed-source
libmali.sofirmware from your device vendor.hwaccel.ml.ymlexpects it at/usr/lib/libmali.so, plus/lib/firmware/mali_csffw.bin; change the paths if yours differ. - Optionally,
MACHINE_LEARNING_ANN_FP16_TURBO=truecan make it much faster at the cost of very slightly lower accuracy.
RKNN
- The vendor’s kernel driver, usually preinstalled.
- RKNPU driver 0.9.8 or later (check with
cat /sys/kernel/debug/rknpu/version). - Optionally, setting
MACHINE_LEARNING_RKNN_THREADSto 2 or 3 can make RK3576 and RK3588 dramatically faster than the default of 1, but multiplies each model’s RAM use by the same amount.
Set it up
Section titled “Set it up”- Download
hwaccel.ml.ymlinto the same folder asdocker-compose.yml. - In the
immich-machine-learningservice, add-cuda,-rocm,-openvino,-armnnor-rknnto the end of theimagetag. - In the same service, uncomment the
extendssection and changecputo your backend:cuda,rocm,openvino,openvino-wsl,armnnorrknn. Use the-wslversion on WSL2 where there is one. - Redeploy the machine learning container.
The result looks like this:
immich-machine-learning: container_name: frameleaf_machine_learning image: ghcr.io/frameleaf/frameleaf-machine-learning:${FRAMELEAF_VERSION:-${IMMICH_VERSION:-release}}-cuda extends: file: hwaccel.ml.yml service: cuda volumes: - model-cache:/cache env_file: - .env restart: alwaysConfirm it’s working
Section titled “Confirm it’s working”- Watch GPU use while a job runs, with
nvtop(NVIDIA or Intel),intel_gpu_top(Intel) orradeontop(AMD). - Or check the machine learning container’s logs (
docker compose logs immich-machine-learning) when a smart search or face detection job starts, or when you search with text. You should seeAvailable ORT providerslisting your provider, such asCUDAExecutionProvider, or for ARM NN aLoaded ANN modelline without errors.
AI descriptions and Locked-content detection
Section titled “AI descriptions and Locked-content detection”Image enrichment uses two separate model paths:
| Task | CUDA | OpenVINO |
|---|---|---|
| Locked-content (NSFW) detection | ONNX Runtime with onnx-community/nsfw_image_detection-ONNX |
ONNX Runtime with OpenVINO |
| Descriptions and tags | Transformers and PyTorch in the CUDA image | OpenVINO GenAI |
The default description model is Qwen/Qwen2.5-VL-3B-Instruct. On Intel integrated graphics it’s mapped to the OpenVINO-converted llmware/qwen2.5-vl-3b-ov. On NVIDIA it runs directly, and the same code path handles larger Qwen2.5-VL models (7B, 32B, 72B) and Qwen3-VL. Larger models need more video memory. The fallback model is microsoft/Florence-2-base-ft.
The CUDA image keeps NVIDIA acceleration for the other models too, and the service still prefers CUDAExecutionProvider when CUDA is installed. See AI descriptions and smart albums.
Model licences
Section titled “Model licences”Some models limit commercial use. Check the licence before using a model outside your own installation.
| Model | Licence |
|---|---|
Qwen/Qwen2.5-VL-3B-Instruct and llmware/qwen2.5-vl-3b-ov |
Qwen Research License Agreement (Alibaba Cloud) |
nllb-clip search models (every variant) |
CC-BY-NC-4.0 |
MusicGen-small (Xenova/musicgen-small) |
CC-BY-NC-4.0 |
Unraid, Portainer and single Compose files
Section titled “Unraid, Portainer and single Compose files”Some platforms can’t use more than one Compose file. Copy the relevant section of hwaccel.ml.yml into the service directly instead of using extends. For CUDA, that’s:
immich-machine-learning: container_name: frameleaf_machine_learning # Note the -cuda at the end, and no extends section image: ghcr.io/frameleaf/frameleaf-machine-learning:${FRAMELEAF_VERSION:-${IMMICH_VERSION:-release}}-cuda deploy: resources: reservations: devices: - driver: nvidia count: 1 capabilities: - gpu volumes: - model-cache:/cache env_file: - .env restart: alwaysThen redeploy the machine learning container.
For video transcoding on these platforms, such as Intel Quick Sync (pass in /dev/dri) or NVENC on Unraid, see Video transcoding.
Several GPUs
Section titled “Several GPUs”For more than one NVIDIA or Intel GPU, set MACHINE_LEARNING_DEVICE_IDS to a comma-separated list of device IDs and MACHINE_LEARNING_WORKERS to how many you listed. Find the IDs with nvidia-smi -L or glxinfo -B.
MACHINE_LEARNING_DEVICE_IDS=0,1MACHINE_LEARNING_WORKERS=2The service then starts two workers, one per GPU, and each request goes to one or the other. To always use one particular GPU, list just that one, for example MACHINE_LEARNING_DEVICE_IDS=1.
Raise the job concurrency to keep several GPUs busy. Each GPU must be able to hold all the models on its own: a model can’t be split across GPUs, and you can’t send one model to one GPU.
- If a model fails, try a different one to see whether the problem is specific to that model.
- Raising concurrency past the default improves GPU use, but also uses more video memory.
- Larger models benefit more from acceleration, if you have the video memory.
- On Rockchip boards, RKNN (the NPU) supports more models than ARM NN, including search, runs cooler and leaves the GPU free for transcoding. It always uses FP16, so it’s very slightly less accurate than ARM NN’s default FP32. With one thread it’s slower than ARM NN for jobs but similar for searching; with three threads it’s somewhat faster than ARM NN at FP32, at the cost of more RAM.