Considered scenarios
Decided 2026-08-20: scenario 3 is retained. Its page is promoted to
Architecture & execution, the executive summary is
the datacenter page, and the two R740xd it buys are ordered. This page and the
scenario pages stay as the decision record; the table below describes the scenarios as
compared — the one post-retention change is that calypsomaster, not rumba, takes the
mgmt-01 watcher role (2026-08-20, for its iDRAC9).
Three scenarios for the same target: ISC³ as a two-site datacenter — production in building 19 (two 42 U water-cooled racks ordered, delivery 30 October 2026, 22 kW chilled water), physical lab and off-site backup in 23N307 (3 kW). All three share the base layer (Proxmox VE everywhere), backups with a copy in the other building, no Ceph in production, and default-drop network zones. They differ in which machine takes which role and what gets bought.
| Scenario 1 | Scenario 2 | Scenario 3 | |
|---|---|---|---|
| What it is | The umbrella plan | Scenario 1 with the storage layer redone | Clean slate — every machine may move and change role |
| Production cluster | epyc0 + epyc1 + rumba (Intel quorum, no migration to it) | same (variant: R7515 as third node) | epyc0 + epyc1 + R7515 + gpu-01 — all AMD, full migration mesh, QDevice on mgmt-01 |
| Primary backup | R7515 as pbs-01 | R7515 as pbs-01 (or → production) | bought Dell R740xd 24 × SFF all-NVMe in 307, ≈ 10.5 TiB on the PM983 surplus — sized on a measured 171 GiB store |
| Second copy | calypsomaster — invalidated by the 2026-08-13 inventory | bought R740xd 12 × LFF, ≈ 70 TiB | none — hardened push token + ZFS snapshots under the datastore; the twin chassis becomes the homes filer · calypsomaster sold |
| Bulk tier | none | bought filer-01, ≈ 70 TiB | 10 × HPE SAS in the R7515 — free, ≈ 9.8 TiB |
| PM983 stock (48) | 12 used, ~24 sold | 14 used | 24 used (12 in pve-03, 8 in pbs-01, 2 in filer-01, 2 in gpu-01), 24 spares |
| Monitoring | VMs on the cluster it monitors | same | rumba as standalone watcher outside the cluster |
| Lab compute | carnaval cluster + pool from decommissioned calypso | same | carnaval grown to 8 nodes (5 taken from the R630 spares) + the 8-node bare-metal pool |
| Student homes | on the FS2500, unchanged | same | filer-01 — a dedicated NFS filer in building 19, twin chassis of pbs-01 (§4) |
| Fast fabric | 100 G back-to-back, pair only | same | CRS520: one 100 G port per production node and for filer-01, L3 hardware offload · S5850 as 10 G aggregation · DACs from stock |
| GPU CI runner | — | — | gpu-01 — G242-Z11 + 3 × RTX PRO 4500 Blackwell, fourth production node; interim on carnaval2's T4 |
| New chassis to buy | 0 (+ UPS, switch, RAM) | 2 (pbs + filer) — ≈ €5 300–7 100 | pbs-01 + filer-01 + gpu-01 + CRS520 — ≈ €22 000 + UPS/PDUs, priced against live ServerShop24 stock |
| Status | Not retained (2026-08-20) — §10–13 stay the common reference | Not retained (2026-08-20) — the 12 × LFF chassis it rests on is not in stock (why) | Retained (2026-08-20) — promoted to Architecture & execution, the two R740xd are ordered |
Scenario 4 is scenario 3 with one delta: the backup growth path is made concrete on hardware in stock today (an HPE D3600 LFF shelf behind the PBS box) instead of the empty MD1200/MD1400 listings. Everything else it inherits from scenario 3 by reference, so it is not a column above.
Facts settled since scenario 1 was written, and true whichever scenario wins: two CCR2216 confirmed
on hand, the FS S3900-48T4S and an FS S5850-32S2Q available, epyc0 available again, and a system
engineer joining end August 2026 — so the "bus factor 1" premise several decisions rested on is now
a two-person team whose knowledge is not yet shared.
The NIC shelf is inventoried: 6 × ConnectX-4 Lx dual 25 G (that card's ceiling — it serves no
100 G port) and at least 6 spare ConnectX-6 VPI, same model as the one fitted in epyc1 —
MT28908/MT4123, single port QSFP56, Gen4 x16, HDR200-capable
(the card). No 100 G card has to be bought in any scenario, but
single-port means one link per card: scenario 1's back-to-back replication DAC needs a second card
per EPYC node, scenario 3 needs one card per node and nothing more. The R7515 shipped on 12 August with the
hybrid backplane — 12 SAS/SATA in bays 0–11, 12 universal NVMe in bays 12–23 — plus two 300 GB 10K
SAS disks that make a SAS rpool free, and two PC4-2666 DIMMs to pull when the 8 × 32 GB arrive.
Common open items regardless of scenario: the building-19 electrical request and fit-out, the SInf uplink move, the inter-building fibre, the R7515 NVMe bay cabling, the UPS purchase — details in scenario 1 §10–13 and scenario 3 §10.