Skip to main content

Considered scenarios

Decided 2026-08-20: scenario 3 is retained. Its page is promoted to Architecture & execution, the executive summary is the datacenter page, and the two R740xd it buys are ordered. This page and the scenario pages stay as the decision record; the table below describes the scenarios as compared — the one post-retention change is that calypsomaster, not rumba, takes the mgmt-01 watcher role (2026-08-20, for its iDRAC9).

Three scenarios for the same target: ISC³ as a two-site datacenter — production in building 19 (two 42 U water-cooled racks ordered, delivery 30 October 2026, 22 kW chilled water), physical lab and off-site backup in 23N307 (3 kW). All three share the base layer (Proxmox VE everywhere), backups with a copy in the other building, no Ceph in production, and default-drop network zones. They differ in which machine takes which role and what gets bought.

Scenario 1Scenario 2Scenario 3
What it isThe umbrella planScenario 1 with the storage layer redoneClean slate — every machine may move and change role
Production clusterepyc0 + epyc1 + rumba (Intel quorum, no migration to it)same (variant: R7515 as third node)epyc0 + epyc1 + R7515 + gpu-01 — all AMD, full migration mesh, QDevice on mgmt-01
Primary backupR7515 as pbs-01R7515 as pbs-01 (or → production)bought Dell R740xd 24 × SFF all-NVMe in 307, ≈ 10.5 TiB on the PM983 surplus — sized on a measured 171 GiB store
Second copycalypsomaster — invalidated by the 2026-08-13 inventorybought R740xd 12 × LFF, ≈ 70 TiBnone — hardened push token + ZFS snapshots under the datastore; the twin chassis becomes the homes filer · calypsomaster sold
Bulk tiernonebought filer-01, ≈ 70 TiB10 × HPE SAS in the R7515 — free, ≈ 9.8 TiB
PM983 stock (48)12 used, ~24 sold14 used24 used (12 in pve-03, 8 in pbs-01, 2 in filer-01, 2 in gpu-01), 24 spares
MonitoringVMs on the cluster it monitorssamerumba as standalone watcher outside the cluster
Lab computecarnaval cluster + pool from decommissioned calypsosamecarnaval grown to 8 nodes (5 taken from the R630 spares) + the 8-node bare-metal pool
Student homeson the FS2500, unchangedsamefiler-01 — a dedicated NFS filer in building 19, twin chassis of pbs-01 (§4)
Fast fabric100 G back-to-back, pair onlysameCRS520: one 100 G port per production node and for filer-01, L3 hardware offload · S5850 as 10 G aggregation · DACs from stock
GPU CI runnergpu-01 — G242-Z11 + 3 × RTX PRO 4500 Blackwell, fourth production node; interim on carnaval2's T4
New chassis to buy0 (+ UPS, switch, RAM)2 (pbs + filer) — ≈ €5 300–7 100pbs-01 + filer-01 + gpu-01 + CRS520 — ≈ €22 000 + UPS/PDUs, priced against live ServerShop24 stock
StatusNot retained (2026-08-20) — §10–13 stay the common referenceNot retained (2026-08-20) — the 12 × LFF chassis it rests on is not in stock (why)Retained (2026-08-20) — promoted to Architecture & execution, the two R740xd are ordered

Scenario 4 is scenario 3 with one delta: the backup growth path is made concrete on hardware in stock today (an HPE D3600 LFF shelf behind the PBS box) instead of the empty MD1200/MD1400 listings. Everything else it inherits from scenario 3 by reference, so it is not a column above.

Facts settled since scenario 1 was written, and true whichever scenario wins: two CCR2216 confirmed on hand, the FS S3900-48T4S and an FS S5850-32S2Q available, epyc0 available again, and a system engineer joining end August 2026 — so the "bus factor 1" premise several decisions rested on is now a two-person team whose knowledge is not yet shared.

The NIC shelf is inventoried: 6 × ConnectX-4 Lx dual 25 G (that card's ceiling — it serves no 100 G port) and at least 6 spare ConnectX-6 VPI, same model as the one fitted in epyc1 — MT28908/MT4123, single port QSFP56, Gen4 x16, HDR200-capable (the card). No 100 G card has to be bought in any scenario, but single-port means one link per card: scenario 1's back-to-back replication DAC needs a second card per EPYC node, scenario 3 needs one card per node and nothing more. The R7515 shipped on 12 August with the hybrid backplane — 12 SAS/SATA in bays 0–11, 12 universal NVMe in bays 12–23 — plus two 300 GB 10K SAS disks that make a SAS rpool free, and two PC4-2666 DIMMs to pull when the 8 × 32 GB arrive.

Common open items regardless of scenario: the building-19 electrical request and fit-out, the SInf uplink move, the inter-building fibre, the R7515 NVMe bay cabling, the UPS purchase — details in scenario 1 §10–13 and scenario 3 §10.