Part 16 of the series: Building a Self-Hosted AI Development Platform
•
5 min read

Building a Self-Hosted AI Development Platform — Part 16: Hardware You Only Think About Once

Part 16 of the Forge series: the firmware settings, kernel workarounds and disk checks that live below every Compose file — including a BIOS bug that flooded the logs once per second

Part 16: Hardware You Only Think About Once

Fifteen parts of this series describe software that a repository can rebuild. Compose files, Caddy routes, CI workflows, backup scripts — all of it is text, tracked, and reproducible.

This post is about the layer underneath, where none of that applies. Firmware settings live in a chip on the motherboard. A kernel boot parameter lives in a file the backup job does not cover. A disk’s health lives in its own controller. You configure these once, forget them for a year, and then discover them again at the worst possible time.

So they get their own document in the repository: what the settings are, why each one is set that way, and what to restore if the firmware ever forgets.

The Baseline Worth Writing Down

The host is the ASUS dual Xeon board from Part 2. Its recorded baseline looks like this:

Item Value
BIOS 3801, dated 2019-08-23
Boot mode Legacy BIOS with CSM required
SATA mode AHCI on both controllers
Virtualization Intel VT-x enabled
Boot disk Samsung 850 PRO, serial recorded
Secondary SSD Samsung 850 PRO, serial recorded, backing ssd-fast

Two things in that table do real work.

The disks are identified by serial number, not by /dev/sda. Device letters are assigned at boot. They move when a controller enumerates differently, or when a disk is added. Part 3 built a ZFS mirror and an LVM thin pool across four disks. The phrase “the boot disk” is only unambiguous once you write down which physical device it means.

Legacy BIOS with CSM is a constraint, not a preference. A board from this era boots this way, and a firmware reset that silently switches to UEFI produces a machine that does not boot at all, with no obvious cause. The document says plainly: if a firmware reset changes the boot settings, restore Legacy/CSM, AHCI on both controllers, and the recorded boot disk before changing anything else.

That sentence exists for a version of me who is already stressed, standing at the machine with a keyboard plugged into it, not thinking clearly.

The Firmware Bug

The most interesting item is a defect in the firmware itself.

The board asserts an ACPI general-purpose event, GPE 0x24, once per second. Its handler references an object that does not exist in the firmware tables, so every assertion produced a kernel error. Once per second, continuously, for as long as the machine was powered on.

Nothing was broken. Guests ran, storage was healthy, ECC counters were clean. But the kernel log filled with the same four error names forever. That makes the log useless, and it is the log you need when something is wrong. It is the inverse of the problem in Part 13. There, nobody read a good signal. Here, a real signal sits buried in noise.

BIOS 3801 is the latest release for this board and does not fix it. The kernel version in use does not work around it. So the workaround is to tell the kernel to ignore that one event.

The test came first, applied live and reversibly:

# Mask only this GPE, and watch what happens
echo "mask" > /sys/firmware/acpi/interrupts/gpe24

The counter stopped, the kernel errors stopped, and guests, storage, ZFS and ECC counters all stayed healthy. Only then was it made persistent:

# /etc/default/grub
GRUB_CMDLINE_LINUX_DEFAULT="quiet acpi_mask_gpe=0x24"

update-grub, then confirm the parameter landed in the generated boot entry. The pre-change file was kept alongside it with a dated suffix, in the same pattern as the network rollback in Part 15: keep the original, name it by date, and the undo step needs no reconstruction.

Why it is not finished

This workaround is recorded as pending, and that distinction is the point of the document.

The live mask proved the fix works on the running system. The GRUB change is supposed to apply it automatically at every future boot — and that has not been proven, because the host has not been rebooted since. A boot parameter that is wrong fails silently: the machine comes up, everything looks normal, and the flood quietly returns.

So the document carries a checklist to run on the next planned reboot, not a claim of success:

# 1. Did the running kernel actually receive the mask?
cat /proc/cmdline        # expect acpi_mask_gpe=0x24

# 2. Is the GPE masked, and is the counter static?
cat /sys/firmware/acpi/interrupts/gpe24
sleep 10
cat /sys/firmware/acpi/interrupts/gpe24

# 3. Did the flood stay away for this entire boot?
journalctl -k -b --no-pager | grep -E 'PRAD|H2RD|HSCI|_L24' || true

Until those pass, the honest status is “works now, unproven across reboots”. That is a different thing from done, and the repository says so.

Disks Earn Their Jobs

The third category is storage hardware, where the platform’s rule is that a disk is tested before it is trusted — especially a disk that already had a previous life.

The Time Machine disk in Part 14 came out of a NAS bay. Before it was allowed to hold three Macs’ backups it had to pass an extended SMART test with zero reallocated sectors, zero pending sectors, zero offline-uncorrectable sectors and zero CRC errors. Its old contents were erased only after confirming the exact drive serial and the existence of two independent, checksum-verified copies of what was on it.

Three details make that safe rather than merely careful. An extended test, because the short test reads a sample and can pass on a failing drive. Zero across all four counters, because a drive with reallocations has started a trend. And matching by serial, because “the 4TB disk” is a description, while a serial is an identity — the failure mode being avoided is erasing the wrong drive.

The Layer Git Cannot Reach

There is a boundary running through this series that this post makes explicit.

Part 9 argued that the repository must be able to rebuild the platform. That is true for everything above the hypervisor: Compose files, configuration, documented decisions. It is not true below it. A git clone does not set a SATA controller to AHCI, does not restore a kernel boot parameter, and cannot tell you which physical disk is safe to erase.

For that layer, documentation is the backup. Not a snapshot that restores automatically, but a written record of what the settings are, which of them are constraints rather than choices, which workarounds are live, and which are still unproven. A rebuild on this hardware starts with that document, and the Compose files only matter after it.

What This Stage Delivered

The physical host now has a recorded baseline, with serials rather than device letters and the boot-mode constraint stated explicitly. A firmware defect has a tested workaround, a kept original file, and a verification checklist that keeps it marked pending until a reboot proves it. Disks have a health gate before they are given data.

Open items are the ones named above: the ACPI workaround awaits its reboot verification, and from Part 15, the host still runs on a single network cable.

Next in the Series

That is the platform down to the metal. The next posts return to the top of the stack, where the services actually earn their keep — and to the question this build keeps circling back to: what a repository can rebuild, what only documentation can carry, and which of the two you reach for first when something breaks.