Skip to main content

I Moved My Home Server From a Xeon Workstation to an Intel NUC in a Day. Here’s What Broke

My old Xeon workstation kept dying with nothing in the logs. How I moved a dozen Docker and Node services to an Intel NUC in one day, the six things that broke, and what I still do not know.

The old box kept switching itself off with nothing in the logs. I never found out why. I moved everything to a mini PC I already owned instead.

The short answer: my home server, a 2012 HP Z620 workstation, died abruptly at least six times in six weeks and left no trace each time. Rather than keep diagnosing or buy a replacement, I moved its dozen services to an idle Intel NUC in one day. The method that made it quick was simple: dump the databases, leave the big data on the NAS, and give the new machine the old one’s address. Six things still broke, and one of the “clever” choices caused a wrong diagnosis five days later.

What you’ll learn

  • How I decided between more diagnosis, buying a replacement and reusing what I had
  • The migration method, and why taking over the old address saves most of the work
  • The six things that broke, with the fix for each
  • What I still do not know, including why the old box was dying

Evidence note: this is my own homelab. The timeline, hardware details and fixes come from the runbook and notes written during the work in August and September 2026, checked against the live server in October. Addresses and names are removed. Where I am estimating or guessing, I say so.

The problem: a server that dies without a word

The server ran my photo library, a monitoring dashboard, a few Node apps, three Postgres databases and the tunnels that make some of them reachable from outside.

On 9 August 2026 it vanished from the network. No ping, no response on any port. Wake-on-LAN did nothing. It needed someone to press the power button.

It happened again on 28 August, and this time I read the machine’s own journal for the boot that died. The last line was a routine web request. After that, nothing: no shutdown sequence, no kernel panic, no out-of-memory kill, no disk error, no thermal event.

Then the gaps got shorter. The journal showed abrupt ends on 28 August (twice), 10 September, 17 September, 19 September and 20 September.

What I could rule out:

Suspect How I checked Result
Disks SMART on both drives Passed, no reallocated or pending sectors
Memory and CPU faults Machine-check and ECC error counters Zero
Load What was running before each death It died at idle, not under load
Mains power Uptime of the network switch it was plugged into The switch stayed up through every event

What was left was the power supply or the board, and I could not tell those apart from software. I never proved the cause. I want to be clear about that, because the rest of this article is about a decision made without a diagnosis.

Advertisement

Two things the outages taught me before I moved anything

Every recovery credential lived on the machine that died. The API key for my network controller and the token for my smart-home system were stored only on that server. With it down, I could not look up which switch port it was on or power-cycle it remotely. Keep the things you need to rescue a machine somewhere other than on that machine.

My monitoring could not see its own host die. The dashboard that watches everything else ran on this box. When the box went, so did the thing that would have told me. The network controller noticed 26 minutes late.

The decision

I had three options.

Keep diagnosing. The next steps were the BIOS event log, the power LED blink codes and a swap-test of the power supply. All of them needed me standing at the machine, and none was certain to find it. Meanwhile the box was dying every couple of days.

Buy a replacement. I set a budget of AED 2,000 and looked at used business mini PCs. A Dell OptiPlex 7070 Micro with a 9th-generation i5, 16 GB and an NVMe drive was listed at about AED 1,250, which would have done the job.

Reuse what I had. An Intel NUC8i7BEH I already owned was sitting idle.

Old: HP Z620 New: Intel NUC8i7BEH
CPU Xeon E5-2665, 8 cores / 16 threads, 2012 Core i7-8559U, 4 cores / 8 threads, 2018
RAM 31 GB ECC 16 GB (32 GB maximum)
System disk A 2 TB hard drive with over 81,000 power-on hours 512 GB NVMe, plus a 2 TB SATA SSD
Graphics A discrete card nothing on the server used Integrated, with Quick Sync
Idle power Roughly 120 to 150 W (my estimate) Roughly 10 W (my estimate)

The NUC has half the threads and half the memory. What made it enough was looking at what the server actually used. Total disk use was 111 GB. Nothing used the graphics card. Memory use sat well under 16 GB. The 16 threads were mostly idle.

So I chose the NUC. It cost nothing, it was available that afternoon, and it removed the nine-year-old hard drive from the picture as a side effect. I have not measured the power draw of either machine, so treat that row as a guess.

What would have changed the decision: if anything on the server had needed the discrete graphics card, or more than 16 GB of memory, I would have bought the OptiPlex and added RAM.

The method

Four rules did most of the work.

1. Dump databases. Never copy their data directories. A Postgres data directory from one version or one machine is not something to trust on another. A dump is.

for c in photos_postgres app1_postgres app2_postgres; do
  docker exec "$c" pg_dumpall -U postgres | gzip > "$c.sql.gz"
done

On the new box I started each stack with an empty database, then restored into it.

2. Leave the big data where it is. The photo library lives on a NAS and the server only mounts it. Nothing needed copying. I kept the same mount point and the same upload path, and the photo server picked up its existing library without re-importing anything.

3. Take over the old address. I built the NUC on a temporary address, then at cutover powered the old box off and gave the NUC its address and hostname. Every tunnel, reverse proxy, bookmark, phone app and script kept working without being edited. This was the best decision of the day, and it comes back later as a problem.

4. Keep the old box intact and off. Rollback was: power off the NUC, power on the old one. Nothing was wiped from it for a week.

The order of restore mattered: network mounts first, then files, then databases, then the apps that depend on them, then scheduled jobs and tunnels last.

What broke

1. A native Node module refused to load. One app failed with ERR_DLOPEN_FAILED on better-sqlite3. The new machine had a newer Node version, and native modules are compiled for a specific one.

sudo apt install -y build-essential
npm rebuild better-sqlite3

2. The NAS would not let the new box mount anything. The NAS exports its shares to one address only: the server’s. While the NUC was on its temporary address, every mount was refused. The mounts could only be tested after cutover, which is the opposite of what you want. If I did this again I would add the temporary address to the NAS’s export list first.

3. Scripts lost their executable bit. I staged the backups through a Windows PC. Windows does not keep Unix permissions, so shell scripts and one compiled binary arrived as plain files and failed with “permission denied”. A chmod +x on each fixed it. Copying Linux to Linux, or using a tar archive end to end, avoids it.

4. One app would not start after a “production” install. The app lives in an npm workspaces repository, so dependencies have to be installed at the repository root, not inside the app’s folder. I also installed with --omit=dev, and the app’s start command runs through tsx, which is a development dependency. Installing without that flag fixed it.

5. Two folders and my SSH keys were never copied. My backup list covered the services I had thought of. It did not include the server’s own ~/.ssh folder or two small folders holding tunnel settings. The keys that let the server log in to other machines were gone, so I generated a new key and authorised it on each remote host, and rebuilt the two tunnel configurations by hand. A full copy of the home directory, minus the large folders, would have been safer than a hand-picked list.

6. A site that had been broken all along. One hostname returned 502 after the move. It turned out it had been returning 502 before the move too. The reverse proxy in front of it was forwarding to port 80, and nothing had ever listened there. I added a small nginx server on port 80 that passes requests to the app. Moving house is when you find out what was already broken.

Verification

  • The photo server reported 54,413 items, the same count as on the old box.
  • Every container and every process-manager app came back on its own after an unattended reboot. That was the property I relied on the old server for.
  • I checked the NUC on 10 October 2026: it had been up for 19 days and 16 hours without a break, since the day of the move. The old box had not managed three days in a row by the end.

The mistake the same address caused

Five days after the move, I asked Claude, the AI assistant I use for this kind of work, to look into the outages again. It got it wrong, and the mistake is worth describing because a person could make it just as easily.

The new machine had the old one’s address and hostname, so nothing on the network said it was a different computer. Claude logged in and found system logs that started on 20 September. It read that as evidence having been wiped. It saw that the crashes stopped right after a kernel upgrade on 20 September and recorded that as the likely fix. It even took the abnormal-shutdown counter from the NUC’s own NVMe drive as proof of the old box’s crashes.

None of that held up. The logs started on 20 September because that was the day the operating system was installed. The crashes stopped because the hardware was replaced. The drive’s counter described the NUC’s earlier life.

Taking over an address is the right call for cutover, and I would do it again. But write the swap down where you will see it, and check the hardware model before you reason about a machine’s history:

sudo dmidecode -s system-product-name

What I still do not know

  • Why the Z620 was dying. The power supply is my best guess. It is unproven. The machine now runs as an occasional desktop, and I have not stress-tested it.
  • Whether the NUC is immune. Nineteen days is not proof. A second NUC of the same family in my rack, with the same 2021 BIOS, switched off without a trace once in September after 82 days of uptime. I have not found the cause of that either.
  • The real power saving. Both figures in the table are estimates. I plan to measure the NUC at the wall.

Checklist

  1. Before anything fails: keep recovery credentials off the server, and monitor the server from somewhere else.
  2. List what the machine actually uses (disk, memory, GPU) before sizing a replacement.
  3. Dump every database. Do not copy data directories.
  4. Copy the whole home directory, including ~/.ssh, not a hand-picked list.
  5. Add the new machine’s temporary address to your NAS export list so you can test mounts early.
  6. Avoid passing Linux files through Windows, or re-apply permissions afterwards.
  7. Restore in order: mounts, files, databases, apps, scheduled jobs, tunnels.
  8. Power the old machine off, then give the new one its address.
  9. Reboot once more and check everything again before trusting it.
  10. Record the hardware change somewhere you will read before the next investigation.
Advertisement

Takeaway

I replaced a server without ever learning what was wrong with it, and that was the right trade. A day of migration bought back a machine I could trust, where another week of diagnosis might have bought nothing. The cost was smaller than I expected and mostly came from my own backup list being shorter than the machine’s real contents.

AI assisted with the migration work and with drafting this article. The servers, the failures and the decisions are from my own homelab.

Written by Engineer running a 24/7 homelab since 2022. Every guide here is built and tested on my own hardware. No paid placements.

Keep reading