Skip to main content

Storage & backup — option 2

Date: August 2026 · Status: not retained — scenario 3 was retained on 2026-08-20. Its pbs + filer pair survives there as two identical R740xd, all-NVMe rather than LFF, sized on the measured store. Scope: the second-site backup copy, and a bulk filer for production. An alternative to §5.2 of the target architecture, which stands as option 1.


1. Why option 1 does not hold

§5.2 puts the second backup copy on calypsomaster "repurposed", sized to match pbs-01's 10.5 TiB backup pool, and sells the surplus PM983 because no working U.2 bay exists for them.

calypsomaster was inventoried over SSH and Redfish on 2026-08-13:

Assumed by §5.2Measured
ModelPowerEdge R740XDPowerEdge R740, tag 8MYS2T2
Baysenough for ≈ 10.5 TiBbackplane BP14G+, 8 × 2.5" SFF, 4 free
Bay typeSAS/SATA, no NVMe support
ControllerPERC H730 Adapter, FW 25.5.9 — hardware RAID
Disks4 × Toshiba THNSN81Q 1.92 TB SATA SSD
CPU / RAM2 × Xeon Gold 6154, 36 c / 72 t, 192 GB

Eight 2.5" bays cannot hold a 10.5 TiB second copy on any sensible drive: 2.5" nearline tops out at 2 TB, so a full chassis is ~16 TB raw and ~11 TB in RAIDZ2, with no room to grow retention. Reaching it on SSD means eight 3.84 TB units, a budget line nobody has costed. §5.4 calls this copy "the layer that does not exist today" — it is the whole off-site leg of 3-2-1, so it should not rest on an assumption that turned out wrong.

Two other gaps, from the same inventory round:

  • Production has no bulk tier. Course datasets, the MAAS image library, ISOs and golden exam images, the GitLab object store and lab scratch have nowhere to live except nvme_pool, which §5.2 sizes at ≈ 5.8 TiB per node for ~4 TB of service load.
  • rumba is at 58 % on hdd (~1.1 TiB usable free) and every DIMM slot is populated, so it is not the answer either.

2. What to buy

The chassis is not currently in stock

The 14 August 2026 market check done for scenario 3 found no Dell 12 × LFF machine listed at ServerShop24 — not the R740xd, not the R730xd, not the R7515 — and no MD1200/MD1400 LFF shelf either. The in-stock Dell LFF machine is an R750xs 8 × LFF at €4 135. The 8 TB SATA line priced below at €95–130 is out of stock as well (€109 when it was listed); the 14 TB LFF is €252, 78 units. Stock moves, so this blocks the purchase rather than the idea — but the §6 total is not quotable today.

There is a shelf, just not a Dell one — scenario 4 specifies it. A shelf covers the filer-01 half of this page, bulk capacity hung off a server that already exists, for a fraction of a new chassis. It does not cover the pbs-02 half, which needs a machine of its own in 307.

Two identical Dell PowerEdge R740XD, 12 × 3.5" LFF, HBA330 in IT mode:

pbs-02 → 23N307filer-01 → building 19
Pool12 × 8 TB SATA, RAIDZ212 × 8 TB SATA, RAIDZ2
Usable≈ 70 TiB≈ 70 TiB
Special vdev2 × PM983, mirrored
CPU / RAMXeon Silver 4210, 64 GBXeon Silver 4210, 128 GB
Network25 G SFP28, mgmt VLAN25 G SFP28 to ccr-core
Rolepull sync from pbs-01datasets, MAAS images, ISOs, GitLab object store, lab scratch

Identical on purpose. The estate is run by a small team (§5.1), so one configuration, one runbook and one spares pool are worth more than fitting each machine to its role. Both are the same iDRAC9/Redfish generation as pbs-01 and rumba, so the existing tooling works unchanged.

12 × 8 TB is well past the 10.5 TiB pbs-02 needs; the headroom is what lets retention grow without a second purchase. 12 × 4 TB (≈ 35 TiB) is the budget floor and still triples the target.

Boot goes on a BOSS card or the two rear 2.5" bays. On filer-01 the rear bays are better spent on the special-vdev mirror, with boot on BOSS.

3. Why LFF, and not a 24-bay NVMe chassis for the PM983

Housing the surplus PM983 needs a 24-bay SFF chassis with the NVMe backplane and PCIe extenders — a specific and uncommon SKU, and the wrong backplane repeats the EPYC situation exactly.

That SKU turned out to be stocked: the R740xd 24 × SFF all-NVMe (RC9) at €1 639, 17 units on 14 August 2026, with 16 × U.2 plus 8 dual-mode bays and no HBA in the path. Scenario 3 builds its whole backup tier on it. What follows still holds as a cost-per-TB argument for a bulk tier; it no longer holds as "the chassis does not exist".

It would also buy little. The drives are Gen3 read-intensive, the workloads are sequential and latency-tolerant, and a 25 G link caps at ~3.1 GB/s — less than two of those drives. Twelve spinners in RAIDZ2 read at roughly 1.5–2 GB/s, and a mirrored NVMe special vdev covers metadata and small blocks, which is what makes an image and dataset workload feel quick. The same link is saturated for a fraction of the money.

The PM983 stock therefore stays on §5.2's plan — 12 into pbs-01, 12 as spares, the rest sold — except that two go into filer-01 as its special vdev.

4. The filer holds reproducible data only

filer-01 at ≈ 70 TiB does not fit in pbs-01's 10.5 TiB backup pool, and never will. The rule has to be written down before the machine exists, or it becomes a primary by accident:

caution
filer-01 is not backed up

It holds data that can be rebuilt — course datasets, MAAS and golden images, ISOs, the GitLab object store, lab scratch. Anything whose loss would matter does not live there. Student homes stay on nvme_pool per §5.3.

The GitLab object store is the one entry that needs a second look: artifacts and the container registry are reproducible, but only if the registry cleanup policy and artifact expiry are set. Without them the bucket becomes the largest thing in the estate.

5. Variant: the R7515 goes to production

pbs-01 is a Dell R7515 with a 64-core EPYC 7702, 24 hybrid bays and an HBA330 already in pass-through. It is the strongest machine in the estate after the two dual-EPYC nodes, and it is assigned to a role that needs bays and almost no CPU.

It is also AMD. §4.1 notes that pve-03 being Intel makes live migration from the EPYC hosts impossible, which is why it carries quorum and non-critical guests only. Making the R7515 pve-03 gives a three-node AMD cluster where live migration works, and moves backup duty onto a chassis chosen for capacity.

Cost: a third R740XD, and rumba leaves production for lab or provisioning duty.

Decide before phase 4. §14 commissions pbs-01 in November–December; building its four pools and then re-roling the machine is the same work twice.

6. Cost

Indicative used-market prices, August 2026 — to be confirmed against actual quotes.

ItemUnitQtyTotal
R740XD 12 × 3.5" LFF, 2 × Silver 4210, 64 GB, HBA330, rails, 2 × PSU€1 400–1 9002€2 800–3 800
8 TB SATA enterprise (recertified)€95–13024€2 300–3 100
25 G SFP28 NIC, if not included€80–1202€160–240
≈ €5 300–7 100

The R7515 variant of §5 adds one more chassis and twelve more drives, roughly €2 700–3 500.

The PM983 sale does not offset this: §5.2 already earmarks it for pbs-01's memory and part of the UPS.

7. What this changes in the umbrella plan

If option 2 is adopted, §5.2 needs four edits:

  • pbs-02 is a new R740XD, not calypsomaster.
  • calypsomaster becomes free for PXE/MAAS and provisioning duty, which §4.3 notes it already ran.
  • filer-01 joins the pool table, with the not-backed-up rule.
  • Two PM983 leave the sale list for filer-01's special vdev.

Independently of this page, §6.1 plans 25 G SFP28 for pve-03 without knowing a ConnectX-6 is already fitted to rumba — and the shelf holds at least six more spare ConnectX-6, so no 100 G card has to be bought for any machine in any scenario. The six ConnectX-4 Lx on hand cap at 25 G and cover the lab side.