Part 15 of the series: Building a Self-Hosted AI Development Platform
•
5 min read

Building a Self-Hosted AI Development Platform — Part 15: The Failover That Failed

Part 15 of the Forge series: a second network card, an active-backup bond, a test that cut the house off from DNS, a byte-for-byte rollback — and a problem that is still open

Part 15: The Failover That Failed

Every other post in this series ends with something working. This one does not.

In September I tried to give the Proxmox host network redundancy: two cables, an active-backup bond, survive one cable or port failing. The configuration applied cleanly. The failover test then cut the house off from DNS. I rolled it back, verified the restoration byte for byte, and the platform has run on a single network cable ever since.

I am writing it up now rather than when it is solved, because the rollback is the part worth copying, and because a build log that only records successes is not a build log.

What I Was Trying to Fix

The host has two Ethernet ports and uses one. Everything on the platform — the Docker host, the DNS resolver, the photo library, the Git server — reaches the network through nic0 into a bridge, into one cable, into one port on one office switch.

Active-backup bonding puts both NICs into a logical interface with one active at a time. If the active path fails, traffic moves to the standby. It is worth being precise about the scope, because bonding gets oversold: this protects against one cable or one switch port failing. It does not protect against losing the switch, and it does not aggregate bandwidth. Active-backup also does not require switch link aggregation, which is why it looked like a safe change to make alone.

The Preparation That Saved Me

Before changing anything I took a baseline, and this is the part to repeat exactly.

Both NICs negotiated 1000 Mbps full duplex once nic1 was brought up. nic0 had zero recorded RX/TX errors. The config includes directory was empty, with no pending interfaces.new file waiting to surprise me. Sixty pings each to the router, the DNS resolver, the Docker host and Proxmox itself came back with zero loss, and a 100-query DNS check passed over both UDP and TCP.

Then the original configuration, the candidate configuration and a rollback script were saved to a dated directory on the host, with a second copy on my Mac. The candidate deliberately preserved the bridge’s existing MAC address, IP address and gateway, so guests saw no change.

ifquery parsed the candidate successfully. The syntax check raised one warning — an existing bridge-fd 0 setting outside its advertised range — which was pre-existing and harmless with STP already disabled. The apply completed with both slaves up and nic0 active.

And a four-minute timer sat on the host ready to restore the original configuration. That matters more than anything else here. Network changes applied over SSH can sever the connection used to undo them. A host-local timer does not depend on the network surviving, so the worst case becomes a four-minute outage rather than a trip to plug in a monitor.

After the apply, with the bond live on the primary path, everything still worked: zero packet loss to all three guests, DNS over UDP and TCP, outbound HTTPS, all 27 containers running, and the public routes answering.

The Test That Broke It

Then the actual point of the exercise: simulate the primary path failing. A second host-local timer, 45 seconds this time, protected the test.

ip link set nic0 down

The bond did exactly what it is designed to do. It selected nic1. From the host’s own perspective everything was fine — root SSH worked, the host reached the router by ping, and the Docker host still had outbound HTTPS.

From my Mac, the DNS resolver had vanished. Ten pings, ten lost. DNS queries timed out over both UDP and TCP. Restoring nic0 brought it back immediately.

So the host survived the failover and its guests did not. That asymmetry is the whole finding. Traffic from the Proxmox host itself traversed the standby path. Traffic to and from a guest on the bridge did not.

The Rollback

The saved rollback service restored the original configuration. A byte comparison against the pre-change copy passed. nic0 is once again the sole physical member of the bridge, nic1 is down, bond0 no longer exists, and the rollback timers are stopped. No reboot was needed, and the household DNS outage lasted as long as the test did.

That is the one unambiguous success in this post. The change was reversible, the reversal was verified rather than assumed, and the evidence lives in the repository.

What I Still Do Not Know

Two days later I identified the office switch model and found a genuine limitation: the manufacturer’s manual explicitly excludes that model from link aggregation support. It is easy to confuse with two near-identically named models that do support it.

That finding is real, and it does not explain the failure by itself. Active-backup bonding does not need the switch to support aggregation — that is the main reason to choose it over LACP in a home setup. So the limitation is a fact about the switch, not yet a cause.

The leading hypothesis is switch-side: MAC learning or per-port VLAN and PVID settings differing between the two ports, so that frames for a guest MAC never move to the second port. The honest position is that this is a hypothesis. Nobody has logged into the switch — it has a management address and a password I do not have — and the two physical port numbers remain unidentified.

One more caveat I want to keep on the record: taking a link down in software is not identical to unplugging a cable. The switch sees a different event. A test that passes one way can still fail the other.

What Has to Happen Next

The next attempt has preconditions rather than a schedule:

  1. Identify the two physical switch ports and authenticate to the switch.
  2. Inspect MAC learning and the VLAN/PVID configuration on both ports.
  3. Re-test with someone physically present, with console access available.
  4. Require that guest DNS passes in both directions on the standby path, and after failing back to the primary.

Until all four hold, the bond stays out. A failover path that has not been proven to carry guest traffic is not redundancy — it is a second way to cause an outage, on a schedule I do not control.

Why Publish This

Three reasons.

The preparation pattern works regardless of the outcome: baseline measurements first, saved original and candidate configs, a parse check, and a host-local rollback timer that does not depend on the network. That pattern turned a failed network change into a short, controlled outage. It is worth copying whether or not your bond works.

The second reason is that the test caught a real defect before it mattered. Configure the bond, watch it apply cleanly, skip the simulated failure, and the platform sits on redundancy that does not work. The discovery then happens during a real cable fault, at the worst possible moment. An untested failover path is not a failover path.

The third is the gaps file argument from Part 9, applied to this blog. A series that reports only the things that worked describes a platform that does not exist. This is what the build log says today: the bond is not in place, the switch ports are unexamined, and the host runs on one cable.

Next in the Series

Part 16 stays with the physical machine. BIOS settings that have to survive a reboot, a firmware quirk that needed a kernel workaround, and the disk checks that happen before hardware is trusted with data — the layer that no Compose file describes, and that nobody thinks about until it fails.

Next: Part 16 — Hardware You Only Think About Once