The usual advice is that Immich’s machine learning takes days on a CPU. On my 2018 mini PC, the search model got through 51,126 items in under an hour.
The short answer: on an Intel NUC8i7BEH with no GPU acceleration, Immich’s search model processed every preview in my library, 51,126 of them, in 57 minutes 39 seconds. The face model ran at 5.2 photos a second on a 3,000-photo sample, which works out to roughly two hours for my 39,546 photos. For a library of this size, a GPU would save an afternoon, once. I would not buy one for Immich.
What you’ll learn
- Real timings for Immich’s two heavy models on an 8th-generation Intel CPU
- How I measured them without touching the library, so you can repeat it
- What the numbers do not cover
- Where the real limits of my setup are, which are not speed
Evidence note: these are measurements from my own Immich server on 10 October 2026. I timed the machine-learning container directly, using the previews Immich had already made. Addresses and paths are removed. One figure is an extrapolation and is marked.
The question
Guides on self-hosting photos, including an older one on this site, tell you that face recognition and smart search are “much faster with a GPU” and that a CPU “takes days for large libraries”. I run Immich on a mini PC with no graphics card to speak of, so I wanted a number instead of an adjective.
Test environment
| Part | Detail |
|---|---|
| Computer | Intel NUC8i7BEH, Core i7-8559U (4 cores, 8 threads, 2018), 16 GB RAM |
| Acceleration | None. The stock machine-learning image, CPU only |
| Immich | 3.3.1, default models: ViT-B-32__openai for search, buffalo_l for faces |
| Storage | Library on a Synology DS723+ (two 8 TB drives, mirrored) over NFS; database on the NUC’s own SSD |
| Library | 54,413 items: 39,564 photos (75 GB) and 14,849 videos (846 GB) |
| Previews available | 51,126 (39,546 photos, 11,580 videos) |
The NUC also runs a dozen other small services. I did not stop them for the test.
Method
Immich’s server does not send originals to the models. It sends the preview image it has already generated for each item. So I did the same thing, from outside Immich:
- Listed every preview file from the database (read-only query).
- Read each one from the NAS over NFS.
- Sent it to the machine-learning container’s own
/predictendpoint, which is the call the server makes. - Timed each read and each model call.
I ran two requests at a time. Immich lets you set how many of each job run at once in its settings, so your own throughput will depend on that number as well as on the CPU. Nothing was written back, so the library and its existing search index were untouched and usable throughout.
Variables I controlled: same machine, same models, both models loaded and warm before timing started, two requests at a time.
Variables I did not: other services on the box, the NAS doing its own work, the network.
Results
Search indexing (CLIP), every preview
| Measure | Result |
|---|---|
| Items processed | 51,126 |
| Errors | 0 |
| Wall-clock time | 57 min 39 s |
| Throughput | 14.8 items per second |
| Model call, median | 0.124 s |
| Model call, 95th percentile | 0.145 s |
| Read from NAS, median | 0.007 s |
| Data read | 12.5 GB |
| Load average during the run | about 3.1 on 8 threads |
Face detection and recognition, 3,000-photo random sample
| Measure | Result |
|---|---|
| Photos processed | 3,000 |
| Errors | 0 |
| Wall-clock time | 9 min 33 s |
| Throughput | 5.2 photos per second |
| Model call, median | 0.280 s |
| Model call, 95th percentile | 0.817 s |
| Faces found | 4,464 |
Extrapolated, not measured: at 5.2 photos a second, my 39,546 photos would take about 2 hours 6 minutes.
Single requests, nothing else running
| Measure | Result |
|---|---|
| Search model, one image | 0.097 s |
| Face model, one image | 0.328 s |
| A text search query | 0.038 s |
| First call after idle (model has to load) | 11.7 s for search, 7.8 s for faces |
What the numbers mean
The CPU is the limit, not the NAS. Reading a preview over the network took 7 milliseconds at the median. The model call took 124. Moving the library to faster storage would change nothing here.
Faces cost about three times as much as search, and vary far more. The slowest 5% of photos took over 0.8 seconds, which I take to be group shots: more faces means more work per image.
Searching is instant. Turning a typed query into something the index can match took 38 milliseconds. A GPU would not make search feel faster.
The one delay you will notice is the first call. Immich unloads its models after a few minutes of inactivity. The first search after a quiet spell waits around 12 seconds for the model to load. That is a memory trade-off, and it is the same with or without a GPU.
For this library, “days on a CPU” was wrong. Search indexing took under an hour. Faces would take about two. Both are one-off jobs that Immich runs in the background while you carry on.
Limitations
- This was not Immich’s own job. I timed the model step and the file read. Immich’s real jobs also write results to the database, and face recognition then groups faces into people. I did not time those.
- Thumbnails and video transcoding are separate and can easily take longer than the machine learning. With 846 GB of video, transcoding is the heavy job on my library, and I did not measure it.
- The face total is an extrapolation from 3,000 photos.
- One machine, one library. A slower CPU, a larger model than the default, or a library ten times the size changes the picture. At 500,000 photos, the same rate means about nine and a half hours for search and over a day for faces.
- Previews, not originals. That is what Immich feeds the models, but it means these figures say nothing about how long it takes to generate those previews in the first place.
Reproduce it
The test needs only what an Immich install already has. In outline:
# 1. list the preview files (read-only)
docker exec immich_postgres psql -U postgres -d immich -At -c \
"select f.path from asset_file f join asset a on a.id = f.\"assetId\"
where f.type = 'preview' and a.\"deletedAt\" is null"
# 2. find the machine-learning container's address
docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' immich_machine_learning
Then post each file to http://<that address>:3003/predict as a multipart form with two fields: image, and entries, a JSON description of the model to run. For search:
{"clip": {"visual": {"modelName": "ViT-B-32__openai"}}}
The paths in the database are as the container sees them, so map the container’s /data to wherever your library is mounted on the host. Run it with two workers and record the time per call.
What I would actually worry about
Speed turned out not to be my problem. These are, and none of them is solved yet:
- It only works at home. I have not exposed it to the internet, because I am wary of opening my home network to reach a photo library.
- Storage is finite. The NAS volume is 7 TB and 4.5 TB is used.
- There is no offsite copy. The drives are mirrored and Immich dumps its database nightly, but a fire or theft takes all of it. This is the gap I should close first.
For context on how I use it: Immich is my archive. My phone still backs up to iCloud, and nothing new has gone into Immich since the import in May. A daily-driver setup with phones uploading constantly would put a steadier, lighter load on the same hardware.
Takeaway
If you are sizing a box for Immich and the library is in the tens of thousands, a recent-ish four-core CPU is enough for the machine learning. Spend the GPU money on a second copy of your photos somewhere else.
AI assisted with running the measurements and drafting this article. The server, the library and every figure are from my own installation.
