Epyc — history & operations
How Epyc got to its current state: the dated operation logs, newest first. Current facts live on the main page; finished work across the whole rack is indexed in the ops journal.
16 DIMMs pulled from epyc0 for the R7515 (2026-08-29)
epyc0 was populated to 1 TB — 32 × 32 GB, two per channel, clocked down to 2933 MHz. The 16
modules in the Dimm0 slots, all Micron 18ASF4G72PZ-3G2E1, were removed for pve-03 (8), for
epyc1's 16 GB pairs (6) and as spares (2). The remaining 16 run at 3200 MHz, one per channel,
512 GiB.
The first POST after the pull reported 480 GiB with CPU1_ChannelL empty — a module that had not
seated, not a dead channel; it enumerated after reseating. Inventory read over Redfish from the
BMC, which serves the last POST's SMBIOS data: it does not change while the machine is off, so a
boot is needed after any DIMM work.
Full hardware checkup on epyc1 (2026-08-13)
A pass over CPU, memory, storage, PCIe, network, thermals and firmware, from the BMC over Redfish and from the host, with no guests on the machine.
Everything outside the memory checked out. Reference figures worth comparing against later:
| CPU | 96 threads, microcode 0xa0011de, 158.6 GB/s aggregate AES-256-GCM; Tctl peaked 44 °C at 20 °C inlet |
| Memory | 0 correctable and 0 uncorrectable ECC events through 167 GiB written; 8 GiB pattern round-trip clean |
nvme_pool | 4.7 GB/s write, 7.8 GB/s read, direct I/O |
| Drives | all six SMART PASSED, no media errors, no reallocated sectors, 0% endurance used |
| PCIe | every link at its full capable speed and width |
| Power | 153 W of 1600 W, four fans at 6300 RPM, every BMC sensor OK |
Channel P: firmware records a map-out, not an absence
The missing 32 GB gained an explanation, though not a fix. Two firmware sources agree that socket P1's channel was deliberately removed from the memory map rather than simply failing to appear — details and the reasoning are on the main page, which owns this diagnosis.
The useful negative result: DRAM Soft Post Package Repair was enabled and changed nothing.
Capacity stayed at 480 GB, both channel-P slots stayed absent, and the BERT record came back
byte-identical on the next boot. That is consistent with the record being stale rather than live —
a channel whose DIMMs return no SPD was never trained, so it cannot be producing a DRAM error now.
The setting was left enabled as a better default for the memory that is present. Modules, seating,
BIOS defaults, the BIOS and BMC updates, and now sPPR are all eliminated; a CMOS/NVRAM clear and a
CPU swap between sockets are what remain (todo).
Found along the way
- A ConnectX-6 nobody had recorded — an unused InfiniBand card on CPU1, described on the main page. It bears on the open HDR100-versus-HDR200 question.
- PSU1 draws no AC by choice — its cord was never plugged, so the BMC reports it
Criticalwhile the FRU still reads. Not a fault; noted on the main page so nobody chases it. - BIOS attributes cannot be written over Redfish on this BMC, so the sPPR change needed the KVM console. The dead ends are listed on the main page.
- A leftover
zfs-import@data_pool.servicefails at every boot, from the pool retired the day before (todo).