IK
← Incident Log
RESOLVED CRITICAL Home Lab Β· incident Β· 08-14-26 Β· overheating, cpu issues, instant shutoff.

Node 2 Overheating, CPU triggering shutdowns.

What happened

On node 2, my CPU seems to be overheating, which causes the whole node to safely shut down. As a result, my node 1 services lose access to media, backup, and the main mass storage unit.

Root cause

On node 2, I decided to implement my secure torrenting solution using OpenVPN. Torrenting with a VPN is CPU intensive because of constant hash decryption, so the CPU uses more power than it would typically need. In a normal world, my 7th-gen i3 on my hp pavilion would be able to handle that, but unfortunately my world is far from normal. For some reason, my OEM CPU fan isn't making contact with the CPU itself, causing the CPU to spike to 100 degrees and triggering the system's safe shutdown to prevent major damage to my hardware.

The fix

I can go two routes. At first, I tried replacing the thermal paste and I ensured my fan was screwed down correctly, but temperatures still spiked abnormally. When I removed the fan, there was no thermal paste residue on it, leading me to believe it was never making contact in the first place.

My Temporary Solution

Unfortunately, I had to take matters into my own hands because I currently don't have the funds to replace node 2, and migrating my services to node 1 to do the heavy lifting takes time.

For now, I placed an old fan by the intake of node 2, and it's blowing cold air into the chassis continuously. With this solution, I'm still seeing abnormal temperature spikes, but it's about 20 degrees cooler and I can still run my torrenting stack. This is a temporary solution until I can implement one of the next permanent solutions.

Solution Option 1

Attempt to replace the CPU fan. I suspect an aftermarket CPU fan will resolve the issue. It's the cheapest, simplest fix I'm willing to gamble on to see if it can save my node.

Solution Option 2

Migrate all torrent services to node 1 and keep node 2 strictly for handling the storage array. No matter what, I need to implement solution 1, but I think this is the best option to ensure my Unraid node stays up 24/7, since uptime is my #1 priority.

As of right now, this is an open case until I implement one of these solutions.

Resolution

I went with solution 1 and purchased a lower cooler for my CPU to make contact. After this, I tested my thermals, and they were significantly lower when the CPU is idle and under higher wattage usage. I also haven't had any issues with the CPU automatically shutting off due to overheating.

Temps during Parity Check

image

// related project

Home Lab β†’

// referenced in

↩ Node 2 Network Speeds Throttled and Reset on Reboot↩ Node 2 Overheating, CPU triggering shutdowns.↩ My favorite Parity / Redundancy System ↩ Linux Container failing after Network Drive SMB unmount