Network Watchdog
A network card can hang while the rest of the server keeps running: the kernel is alive, the disks are busy, and nothing answers from outside. Without help that server stays dark until someone power-cycles it. One Intel I219 card hang on a Proxmox VE host left the machine unreachable for 23 hours. With a network watchdog, the same fault ends in an automatic reboot 30 to 90 minutes after the network is gone. Pulsed Media ships one in PM Software Stack (PMSS), and this page shows how it decides, how to run the same check on a Proxmox host, and what the two NIC families we have seen hang actually log.
What a NIC hang looks like
The machine is up. Uptime keeps climbing, cron jobs run, the journal keeps writing. From outside there is no ping, no SSH and no ARP reply. Nothing on the box notices, because nothing on the box is broken except the transmit path of one network card.
That is why a hang like this outlasts ordinary monitoring: an external alert tells you the host is gone, but somebody still has to walk to it, or reach it over IPMI, a KVM or a power relay. Consumer boards such as the ASRock DeskMini A300 / X300 have no IPMI at all. The cheapest fix is to let the machine notice by itself and reboot.
How the check decides
The PMSS check is a short shell script, published in the PMSS repository as etc/seedbox/config/template.watchdog.network-check.sh. On every run it:
- pings each IPv4 default gateway, then 1.1.1.1, then 8.8.8.8 (two packets, five-second timeout each);
- exits 0 and clears its failure timer as soon as any target answers;
- on the first failed run writes the current time to
/run/pmss-watchdog-net-failand exits 245 ("loss seen, not yet failing"); - exits 1 only when all targets have been unreachable for 1800 seconds.
The 30-minute, all-targets rule is deliberately lenient. A reboot of a shared seedbox host interrupts every user on it, so a lost gateway, a provider routing blip or one dead upstream resolver must never trigger it. The cost of that leniency is real: a card that is only partly stuck and still answers the occasional ping keeps resetting the timer. We have seen that stretch one Realtek outage past five hours. Do not tighten the threshold to fix that; a false reboot on a multi-tenant host costs more than a slow true one.
The failure timer lives in /run, which is cleared at boot, so every boot starts the 30 minutes from zero.
On a PMSS server
PMSS installs the Debian watchdog daemon with /etc/watchdog.conf and the check above as /etc/watchdog.d/network-check.sh. The daemon runs the check every 10 seconds. When the check returns 1, the daemon keeps retrying for retry-timeout = 3600 seconds before it shuts the system down cleanly. End to end, a NIC hang on a PMSS host recovers in roughly 90 minutes: 30 minutes for the check, then the daemon's retry window. PMSS enables the daemon only when the host has a watchdog character device and leaves it alone if an operator has masked the unit.
Check that it is active:
systemctl is-active watchdog ls -l /etc/watchdog.d/network-check.sh journalctl -b -1 -u watchdog | tail
The last command shows the previous boot's watchdog log. After a recovery it contains lines like test binary /etc/watchdog.d/network-check.sh returned 1, then Retry timed-out at 3620 seconds and shutting down the system because of error 1.
The watchdog matters more than any single NIC fix. We have had a host with the unit disabled stay unreachable for days on a Realtek stall; with the unit enabled, the same fault class recovered without anyone touching the machine in about 90 minutes.
On a Proxmox VE host
Do not install the Debian watchdog package on a Proxmox node. Proxmox's own watchdog-mux service, used by HA fencing, "keeps the underlying /dev/watchdog device open for its entire lifetime, even when no HA client is connected" (Proxmox HA documentation). A second watchdog daemon would compete for the same device.
The check itself does not need a watchdog device. Run it from a systemd timer and reboot when it returns 1. This is what we run on our own Proxmox host:
1. The check. Copy template.watchdog.network-check.sh from the PMSS repository, unchanged, to /usr/local/sbin/pmss-network-check.sh and make it executable.
2. The wrapper, /usr/local/sbin/pm-net-watchdog.sh:
#!/bin/sh # Reboot only when the check reports sustained loss of every target. # Exit 245 = loss seen but under threshold, 0 = reachable: do nothing. /usr/local/sbin/pmss-network-check.sh [ "$?" -eq 1 ] || exit 0 logger -t pm-net-watchdog "all network targets unreachable for 30 min; rebooting" systemctl reboot
3. The service, /etc/systemd/system/pm-net-watchdog.service:
[Unit] Description=Network watchdog check (reboot after sustained loss of all targets) After=network-online.target [Service] Type=oneshot ExecStart=/usr/local/sbin/pm-net-watchdog.sh
4. The timer, /etc/systemd/system/pm-net-watchdog.timer:
[Unit] Description=Run the network watchdog check every minute [Timer] OnBootSec=5min OnUnitActiveSec=1min AccuracySec=10s [Install] WantedBy=timers.target
Enable and test it:
chmod 0755 /usr/local/sbin/pmss-network-check.sh /usr/local/sbin/pm-net-watchdog.sh systemctl daemon-reload systemctl enable --now pm-net-watchdog.timer systemctl start pm-net-watchdog.service systemctl show -p Result,ExecMainStatus pm-net-watchdog.service
A healthy host shows Result=success and ExecMainStatus=0. To test the reboot decision without rebooting, copy the wrapper, point it at a stub check that exits 1, and put a fake systemctl that only echoes its arguments first in PATH. Only exit 1 should print systemctl reboot.
Because there is no extra retry window, this setup reboots about 30 minutes after the last successful ping. Proxmox stops the guests before it reboots, so the reboot takes as long as your slowest guest shutdown. To remove it: systemctl disable --now pm-net-watchdog.timer, then delete the two units and the two scripts.
The same recipe works on any systemd host without a usable watchdog device.
Which watchdog for which failure
| Failure | Kernel alive? | What catches it |
|---|---|---|
| NIC transmit hang | Yes | Network check (this page) |
| Kernel hang or lockup | No | Hardware watchdog device (for example iTCO_wdt), fed by a daemon or, on Proxmox, by watchdog-mux
|
| HA node lost from the cluster | Varies | Proxmox HA self-fencing via watchdog-mux
|
A userspace check cannot reboot a machine whose kernel has stopped. If kernel hangs are your problem too, configure a hardware watchdog module; on Proxmox that goes in /etc/default/pve-ha-manager as WATCHDOG_MODULE=, per the same Proxmox documentation.
Example: Intel e1000e "Detected Hardware Unit Hang"
The e1000e driver covers the Intel onboard gigabit controllers found on many desktop and small-server boards, including the I217, I218 and I219. Its signature:
e1000e 0000:00:1f.6 eth0: Detected Hardware Unit Hang: TDH <61> TDT <92> next_to_use <92> next_to_clean <60>
What it means. TDH is the transmit descriptor head, the slot the hardware is working on. TDT is the tail, the slot the driver has queued up to. In the kernel source (drivers/net/ethernet/intel/e1000e/netdev.c, Linux 6.8) the driver flags a hang when a queued packet has waited past its timeout and the card is not paused by flow control. It re-checks once after flushing descriptor write-backs, and ignores the event as a false hang if head and tail are equal. A head stuck behind the tail means the card has stopped sending.
What we saw. On an Intel I219-LM in a Proxmox host the line repeated every two seconds for 23 hours, 42,108 times in one boot, with TDH and TDT frozen at the same values throughout. The driver never logged Reset adapter unexpectedly, its own recovery path, so it never recovered by itself. The host's journal kept writing the whole time. A power cycle brought it back.
Diagnose after the fact:
journalctl -b -1 -k | grep -c "Unit Hang" journalctl -b -1 -k | grep -m1 "Unit Hang"
A count in the thousands with the journal running to the moment of the reset means a NIC hang, not a crash or a power loss.
Mitigation. The workaround we applied is disabling TCP and generic segmentation offload:
ethtool -K eth0 tso off gso off
To keep it across reboots, add post-up /usr/sbin/ethtool -K eth0 tso off gso off under the interface stanza in /etc/network/interfaces. Treat it as an experiment: we have not yet seen whether it prevents the next hang, and the underlying problem sits in the card or its driver. The watchdog is what bounds the damage if it hangs again. If it keeps hanging with offloads off, the durable fix is a different network card.
Example: Realtek RTL8111H on DeskMini-class boards
The ASRock DeskMini A300 / X300 uses a Realtek RTL8111H, the chip the in-kernel r8169 driver calls RTL8168h/8111h. Its signature is the kernel's generic transmit watchdog firing, followed by a driver reset that does not complete:
r8169 <pci> eth0: NETDEV WATCHDOG: CPU: 7: transmit queue 0 timed out 5020 ms r8169 <pci> eth0: rtl_rxtx_empty_cond == 0 (loop: 42, delay: 100)
What it means. The first line comes from the kernel's network scheduler (net/sched/sch_generic.c): a transmit queue made no progress for five seconds, so the kernel calls the driver's timeout handler. In r8169 that handler schedules a chip reset, and the reset waits up to 42 × 100 microseconds for the transmit and receive FIFOs to drain. The second line says they never drained. The card is stuck, the host is unreachable on every protocol, and the kernel carries on normally.
Drivers. Two drivers exist for this chip: r8169, which ships in the kernel, and r8168, Realtek's out-of-tree driver installed as r8168-dkms. The DeskMini page covers the history. Our later experience does not support its r8168 advice. Hosts on r8168-dkms stall the same way. One host that moved to r8169 together with a 6.12 kernel shows no stalls in the logs we keep. Another, on r8169 with no r8168 installed, kept stalling on the 6.1 kernel and again after moving to 6.12. No driver or kernel change has proven to end these stalls. We prefer r8169 because it ships with the kernel, so a kernel upgrade never waits on a DKMS rebuild.
Firmware. r8169 requests the file rtl_nic/rtl8168h-2.fw for this chip and carries on without it if the file is missing:
r8169 <pci>: firmware: failed to load rtl_nic/rtl8168h-2.fw (-2)
On Debian the file comes from the firmware-realtek package in the non-free-firmware component (non-free on Debian 11), so that component has to be enabled in your APT sources. Missing firmware is only a candidate cause of these stalls: we have seen it together with the stall on the same boot, nothing more. Installing it is cheap and easy to undo. The driver requests the file when the interface comes up, so it takes effect at the next boot:
apt install firmware-realtek
The diagnosis after the fact is the same as for e1000e, with a different pattern:
journalctl -b -1 -k | grep -c "NETDEV WATCHDOG"
Limits
- Dead uplink. If the cable, switch port or upstream is down for good, the host keeps rebooting until it comes back: about every 35 minutes with the timer, about every 90 minutes under the PMSS daemon. That is the price of recovering unattended.
- Partial hangs. A card that still answers an occasional ping keeps the timer from expiring, so recovery can take hours instead of minutes.
- Kernel hangs. A userspace check cannot help when the kernel itself stops; that needs a hardware watchdog.
- Root cause. The watchdog bounds the outage. It does not fix the card. Read the previous boot's kernel log after every watchdog reboot.
FAQ
Will a network watchdog reboot my server during a provider outage? Only if the default gateway, 1.1.1.1 and 8.8.8.8 all stop answering for 30 minutes in a row. An outage that long has already cut every user off, so the reboot costs little.
Why reboot instead of reloading the NIC driver? A reboot clears the hang whatever the card, and the watchdog needs no knowledge of the driver. On a bridged Proxmox host, unloading the driver also removes the port from the bridge, which makes a driver reload fragile.
How do I know a reboot was the watchdog?
journalctl -b -1 -u watchdog on a PMSS host, or journalctl -b -1 -t pm-net-watchdog for the timer setup, shows the watchdog's decision.
Does every Pulsed Media seedbox have this? Not every host: PMSS enables it only on hosts that have a watchdog device. Where it runs, shared seedbox users do not need to configure anything.
See also
- PM Software Stack — the platform that ships the network check
- Installing PM Software Stack — run PMSS on your own server
- ASRock DeskMini A300 / X300 — the board behind the Realtek example
- Network Diagnostics — reading network faults on a seedbox
- Proxmox VE — the hypervisor in the timer example
PMSS is open source under GPL-3.0, so the watchdog above runs on any Debian server with a watchdog device that you install it on, and the timer recipe covers servers without one. Install it yourself, or let a Pulsed Media seedbox run it for you on our hardware.