ISC³ — Architecture & execution
Date: August 2026 · Status: retained (2026-08-20) — the plan being executed. The two
R740xd (pbs-01 + filer-01) were delivered in the last week of August 2026. The executive summary is
the datacenter page; the alternatives that were weighed are under
considered scenarios. This page was scenario 3 ("clean slate"),
promoted on retention; the one change since is calypsomaster taking the mgmt-01 role from
rumba (2026-08-20).
Scope: the whole fleet. Two sites: building 19 (two 42 U water-cooled racks ordered, delivery
30 October 2026, 22 kW chilled water) and 23N307 (existing rack, 3 kW).
Block diagram
1. Role map — every asset
| Asset | Becomes | Site |
|---|---|---|
epyc0 — Gigabyte R282-Z92, 2 × EPYC 7443, 512 GB | pve-01, production cluster | 19, rack-A |
epyc1 — Gigabyte R282-Z92, 2 × EPYC 7443, 480 GB | pve-02, production cluster | 19, rack-A |
| R282-Z92 spare — bought new, August 2026 · 2 × EPYC 7302 (Rome), 24 × U.2, no RAM, no drives | cold spare for the pve-01/pve-02 pair, unracked · carries the estate's only untested 24 × U.2 backplane (todo) | — |
Dell R7515 (epyc3) — EPYC 7702 64 c, 24 hybrid bays | pve-03, production cluster + bulk filer | 19, rack-A |
calypsomaster — Dell R740, 2 × Gold 6154, 192 GB, iDRAC9 | mgmt-01 — standalone watcher: PDM, NOC, SOC, MAAS, QDevice. Free today — nothing load-bearing runs on it (inventory) | 19, rack-A |
rumba — Dell Precision 7920, 2 × Gold 6140, 192 GB | lab / provisioning duty once its guests migrate to the cluster — exact use to define (todo); its ConnectX-6 returns to the shelf | — |
carnaval0-2 — 3 × R630, 2 × A2 + 1 × T4 | carnaval0…2, the GPU nodes of the 8-node lab cluster | 19, rack-B |
| 5 of the 7 × R630 spares | carnaval3…7 — CPU-only lab nodes, same cluster | 19, rack-B |
| the 2 remaining R630 spares | cold spares / parts depot; the count is reconciled at eighteen (register) — surplus can be sold | 19, rack-B |
calypso0-7 — 8 × R630 | pool-0…7 — bare-metal pool, MAAS, off by default. Postponed (2026-08-28): these eight are carnaval nodes for the 2026/27 year, the pool comes after the move | 307 |
tango0 / tango1 — ThinkStation PGX | AI pair, unchanged | 307 |
nas — Synology FS2500 | cold tier / archive — monthly export of configs and homes (§4) | 307 |
| 2 × CCR2216 | ccr-core (one in service) + a cold spare in the rack, unpowered | 19, rack-A |
| CCR2004 | ccr-edge | 307 |
| FS S3900-48T4S | sw-access — 1 G + corosync ring0 | 19, rack-A |
| FS S5850-32S2Q | sw-10g — 32 × 10 G SFP+, 2 × 40 G uplink to core | 19, rack-B |
| FS S3600-48T4S | sw-307 — lab access | 307 |
| CRS326-24G-2S+RM | sw-oob — the OOB island is pure L2 at 1 G, so its single 800 MHz core and missing L3 offload do not matter there | 19, rack-A |
| ConnectX-4 Lx dual 25 G (low profile, Gen3 x8), more than the design needs (stock recount 2026-08-30) | all 10 10/25 G links — mgmt-01, the lab nodes and pbs-01. 25 G is the Lx ceiling, so none of these serves a 100 G port | both |
≥ 6 × ConnectX-6 VPI spare — MT28908/MT4123, single port QSFP56, Gen4 x16, HDR200-capable — plus the cards fitted in epyc1 and rumba | every 100 G port on sw-100g: pve-03, gpu-01, filer-01, and pve-01 if epyc0 carries none. One port per node is all this design needs, so single-port costs nothing here. rumba's card joins the spares when it leaves production | 19 |
| DAC stock, 10 G → 200 G HDR | every link in the design, both sites — nothing to buy | both |
| BlueWalker 2 kVA UPS · APC AP8681 PDU | stay in 307 — UPS feeds pbs-01/nas/edge, PDU enforces the pool duty cycle | 307 |
Disk stock:
| Stock | Goes to |
|---|---|
| 8 × PM1735 3.2 TB (Gen4 AIC, 3 DWPD) | 4 per EPYC node — no spare on the market today, see the buy list |
| 54 × PM983 1.92 TB (U.2 Gen3), 6 of them bought with the R740xd | 12 in pve-03's universal bays 12–23, the only working U.2 housing in the estate (EPYC backplanes are dead) · 8 in pbs-01, RAIDZ2 · 2 as filer-01's homes mirror · 2 as gpu-01's mirror · the R740xd boot from their BOSS M.2 pairs, not from PM983 (the Toshiba SSDs stay in calypsomaster) · 30 spares, sell what is still idle after a year |
| 16 × HPE SAS 1.8 TB 10K | 10 in pve-03 bulk · 2 as its rpool mirror · 4 shelf spares |
| 2 × 300 GB 10K SAS, on the R7515 order sheet | not delivered — rpool runs on two of the HPE 1.8 TB instead |
4 × Toshiba 1.92 TB SATA SSD (in calypsomaster) | stay — mgmt-01's local pool, H730 in HBA mode first (conversion notes) |
8 × 500 GB SATA HDD (in rumba) | follow the machine to its lab/provisioning duty |
2. Decisions
- Two people, from end August 2026. A system engineer joins to build ISC³. The estate is not yet shared knowledge, so the design keeps optimising for a small number of distinct things to know — one base layer, one link speed, one card model, one runbook per role. The window between the arrival and the racks landing on 30 October is when the physical-inspection items in §10 get cleared, which is what turns the buy list from a range into a number.
- One base layer: Proxmox VE everywhere. Re-purposing a machine = destroy/redeploy guests, never a bare-metal reinstall.
- All-AMD production cluster, four nodes —
epyc0,epyc1,epyc3(the R7515) andgpu-01. Guests usecpu: x86-64-v3, which Zen 2 and Zen 3 both provide; any CPU-only guest live-migrates to any node. Sizing rule unchanged: production fits on two nodes — those two plus the backup pair are the production core, the other two are extensions. - The GPU node is a production node, not a lab node.
gpu-01is bought for the GitLab CI runners, and its CPU, RAM and NVMe are production capacity the rest of the time — so it joins the cluster in rack-A rather than sitting apart in the lab rack: one management plane, PBS backups, the same firewall zones. Two consequences: a guest holding a passed-through GPU can neither live-migrate nor be HA-restarted (no other node has a card), so those VMs are pinned and excluded from HA groups; and four nodes is an even vote count, somgmt-01runs a corosync QDevice as fifth vote and the cluster survives losing two nodes. The runner VMs execute student code by design, so they sit inlab-virtuallike any student guest — the chassis is production capacity, its runners are not trusted. - Monitoring outside the cluster.
mgmt-01runs PDM, NOC (Prometheus/Grafana/Alertmanager/Uptime Kuma), SOC (Wazuh/Loki), MAAS. When the cluster is down, the machine that says so is up. The machine iscalypsomaster(decided 2026-08-20, replacingrumbain the role): it has an iDRAC9, so the one box that must be up when the cluster is not is recoverable from the OOB island —rumbais a workstation with no BMC; its Gold 6154 are the faster CPUs, its RAM can still grow (rumbais maxed), and it is free today whilerumbahosts production until phase 2.rumbamoves to lab/provisioning duty once emptied. - ZFS +
pvesrreplication, no Ceph. Asymmetric nodes, a 5–15 min RPO that is acceptable here, and a two-person team whose knowledge is not yet shared. A degraded Ceph cluster asks for decisions across every guest at once; a ZFS pool fails one node at a time and says so in one command. Ceph is a course on the pool, not production. - The lab cluster is physically separate from production.
carnaval0…7in rack-B: templates, quotas, no HA, no backup, disposable. Students cannot starve GitLab. - Eight lab nodes, five of them taken from the R630 spares. The three GPU nodes were never the constraint — CPU-only labs (Docker, Kubernetes, Ansible, networking) fill them first and leave the cards idle. Five spares racked and powered turn a parts depot into ~1 TB of RAM and ~300 threads of lab capacity, and the remaining two stay cold as the parts depot they already are. The GPUs stay where they are: 3 of the 8 nodes have one, and a lab that needs CUDA gets one of those nodes.
- Bare-metal pool off by default. 8 ex-calypso R630 in 307 on switched outlets, MAAS-provisioned per course, wiped after. Power is the budget: measure R630 draw at the PDU before fixing how many run at once (provisionally 4).
- One backup box, in the other building.
pbs-01— a Dell R740xd 24 × SFF all-NVMe in 307 — takes a nightly vzdump of every guest over the fibre, plus the homes fromfiler-01. Two identical boxes with a pull sync were the earlier draft; the store measures 171 GiB, so the second chassis becomes the filer instead. The two are still bought together — one spares pool, and each is the other's cold-spare chassis, both roles rebuilt from a playbook. What the pull bought ("a compromised source cannot delete its copy") is replaced at zero cost: the token PVE holds isDatastoreBackup-only (no prune, no delete), the datastore sits on ZFS with its own snapshots, and the monthly FS2500 export stays. Restore test each semester. - The backup store is all-flash, out of the PM983 surplus — no disk bought. The store that holds every
rumbaguest today, deduplicated, is 171 GiB (measured 2026-08-14 onsrv-pbs, 11 guests). Eight PM983 in RAIDZ2 give ≈ 10.5 TiB usable, sixty times that, and the design's future estate — homes, GitLab, the Calypso services — adds a couple of TiB of primary data, three to six once retention is applied. Flash removes the special vdev and its "mirror it or lose the pool" trap, makes the weekly verify a job of minutes instead of hours, and leaves 14 bays free. - 100 G for every production node and for the homes filer, routed in hardware. A CRS520-4XS-16XQ (16 × QSFP28 + 4 × SFP28, ~CHF 1 700) is the fabric switch:
pve-01/02/03,gpu-01andfiler-01each on their own QSFP28 port,mgmt-01on an SFP28, one 100 G uplink toccr-core. Its L3 hardware offload routes the fat internal flows (replication, migration, NFS, mgmt↔storage) at line rate — the CCR2216 CPU only sees what needs the stateful inter-zone firewall. Seven of the 16 QSFP28 ports are used, port by port in §6; 9 stay free. - One link speed, one card model, one cable type.
filer-01at 100 G is not about throughput — 71 homes of small-file NFS sit far below what 25 G already carries. It is about a fleet where every server link is the same thing, so there is one spare card on the shelf and one procedure to replace it. Ports are not the scarce resource (9 free), and cables are not either: the DAC stock on hand runs 10 G to 200 G HDR. What this costs is a 100 G card forgpu-01andfiler-01if the ConnectX-4/5 stock turns out to hold none — €212–283 each. - No 200 G fabric. No MikroTik 200 G switch exists; the alternatives (Mellanox SN3700 class) cost, draw and blow too much for pools that measure 4.7–7.8 GB/s (≈ 40–60 Gb/s) — the disks bound the flows, not 100 G. 200 G stays where it is free: point-to-point DACs (the
tangopattern), one cable if a pair ever needs it. - GPU CI runners on a dedicated GPU node.
gpu-01— Gigabyte G242-Z11 (4 × PCIe Gen4 x16) with 3 × RTX PRO 4500 Blackwell 32 GB, Server Edition (dual-slot, passive, 165 W, 12VHPWR). Three dual-slot cards leave one Gen4 x16 free for the 100 G NIC, so the card does not depend on the riser. One GPU passed through whole perrunner-gpu-0xVM: gitlab-runner, docker executor, NVIDIA toolkit, taggpu; the third card serves course inference and is the warm spare — nothing else in the estate takes an RTX PRO 4500's place, so it moves into a runner VM if one fails and inference stops instead of CI. Runners are rebuilt from a playbook, not backed up. Interim until it lands:carnaval2's T4 passed through to a first runner VM. A2/T4 go back to student CUDA afterwards. - Every server ≥ 10 G. The ConnectX-4/5 stock retires the 1 GbE data path everywhere, lab nodes included.
- The homes get a dedicated filer.
filer-01— the twin of thepbs-01chassis, standalone PVE host in rack-A, 100 G, homes on a PM983 mirror (§4). It takes both halves of the trade the earlier options split: thesec=sysblast radius stays a box that holds nothing but homes, and the mechanics are ZFS end to end — quotas, snapshots, the PBS chain. The cost is one more machine and no HA; accepted, because the mount is non-fatal and the chassis is interchangeable withpbs-01's. bulkholds reproducible data only (ISOs, MAAS images, datasets, scratch) — not backed up, written down before it exists.- OOB island. Every BMC/iDRAC on
sw-oob, no route except the admin VPN group.
The production core
What has to work is small: the reverse proxy and the sites behind it, GitLab CE, Keycloak, the
mail relay, the DBs. The 11 guests running on rumba today measure 171 GiB deduplicated in PBS
(§5); GitLab is the only sizeable addition. Moodle is not in the list —
ISC Learn goes to managed hosting, off the rack, so it survives a site
loss. The rest of this scenario is sized for the lab estate, not for these services.
The core is the subset those services need, and it stands alone:
pve-01+pve-02— ZFS +pvesr, the QDevice onmgmt-01as third votepbs-01in 307 — nightly, restore test each semestermgmt-01watching from outside the cluster- the zones minus the two lab ones —
lab-virtualandlab-metalarrive with the lab - the existing switches — no core flow exceeds 10 G
Everything else is an extension: still ordered in phase 0 (lead times), but adopted one at a time, and the core never waits on one. When an extension slips, fails or is being debugged, production does not notice.
| Extension | Serves | Adopted when |
|---|---|---|
pve-03 as third node | capacity, bulk | its backplane cabling is verified (§10) |
sw-100g (CRS520) | replication/migration speed, one link speed fleet-wide | the offload/routing question (§10) is settled — until then the core runs on the existing 10 G, same VLANs, slower ports |
gpu-01 + CI runners | GPU CI, course inference | delivery (phase 3); runners zoned lab-virtual, never srv-internal |
Homes off the FS2500 onto filer-01 (§4) | ZFS quotas, snapshots, the PBS chain | phase 4 — the FS2500 export stays read-only for a semester as the fallback |
Bare-metal pool, carnaval growth | courses | Calypso wind-down and the rack move — lab side, never in the core's path |
3. Compute
| Node | Machine | CPU | RAM | Pools |
|---|---|---|---|---|
pve-01 | R282-Z92 (epyc0) | 2 × EPYC 7443, 48 c | 512 GB | rpool mirror · nvme_pool 4 × PM1735, 2 mirror vdevs ≈ 5.8 TiB |
pve-02 | R282-Z92 (epyc1) | 2 × EPYC 7443, 48 c | 480 → 512 GB | same |
pve-03 | Dell R7515 (epyc3) | EPYC 7702, 64 c | 256 GB | rpool 2 × HPE 1.8 TB SAS mirror, bays 0–1 · bulk 10 × SAS, 2 × RAIDZ2(5) ≈ 9.8 TiB, bays 2–11 · nvme_pool 12 × PM983, 6 mirror vdevs ≈ 10.5 TiB, universal bays 12–23 |
gpu-01 — to buy | Gigabyte G242-Z11 | EPYC 7443, 24 c — 2.85 GHz base / 4.0 GHz boost, 128 MB L3 · not a P SKU, so one spare chip fits it, pve-01 and pve-02 | 256 GB | nvme_pool 2 × PM983 mirror in the rear bays, from stock + 3 × RTX PRO 4500 Blackwell 32 GB |
carnaval0…7 | 8 × R630 | 2 × E5 v3/v4 | 112–128 GB each | node-local fast ZFS pool, single NVMe/SSD |
mgmt-01 | Dell R740 (calypsomaster) | 2 × Gold 6154, 36 c | 192 GB | rpool + local pool on the 4 × Toshiba 1.92 TB SATA SSD — H730 in HBA mode first |
filer-01 — delivered | Dell R740xd 24 × SFF all-NVMe | 2 × Gold 6244, 16 c | 32 GB | rpool on the BOSS 2 × 240 GB M.2 mirror · homes 2 × PM983 mirror ≈ 1.7 TiB, 22 bays free |
Prerequisites:
pve-02: fourth PM1735, rebuildnvme_poolas two mirror vdevs; fix the dead memory channel (CMOS clear, then CPU swap).pve-03: checked on the bench 2026-08-29. The hybrid backplane is cabled — the machine enumerates 12 U.2 slots at x4, sonvme_poolhas its housing — and the chassis carries four PCIe x16 slots plus an x8, not the two the order sheet suggested, so the NIC competes with nothing. 256 GB fitted from theepyc0DIMMs, 8 × 32 GB at 3200 MHz, one per channel. HBA330 in HBA mode, both 1600 W PSUs, iDRAC9 Enterprise licensed. The two 300 GB SAS the order promised were not delivered:rpooltakes two of the 1.8 TB instead. Takes a ConnectX-6 from the spares.gpu-01: three dual-slot GPUs take six of the chassis' eight slot widths, leaving one Gen4 x16 for the ConnectX-6 — the CRSG120 riser is not needed. The chassis ships 2 × 16 GB, so 256 GB across eight balanced channels means 8 × 32 GB and the shipped pair goes to the shelf.carnaval3…7: a 10 G NIC each from the CX-4/5 stock, a PERC battery where one is missing, and a second PSU where the machine has one slot populated — same R630 hardware order as the existing three.mgmt-01: rescue the residual data oncalypsomaster, put the H730 in HBA mode, reinstall PVE over the iDRAC, static IP in the answer file — the step-by-step conversion plan is in ops todo → calypsomaster.- NICs:
pve-01/pve-02keep their ConnectX-6 (Ethernet mode, 100 G →sw-100g;epyc0's card to confirm);pve-03,gpu-01andfiler-01take a ConnectX-6 from the spares;mgmt-01, the lab nodes andpbs-01take the ConnectX-4 Lx, of which there are 6 for 10 links. - All nodes: populate both PSUs; plug both cords.
- Same storage ID
nvme_poolon all four production nodes, or guest configs need editing at each migration. - ZFS defaults everywhere: mirrors for VM workloads,
compression=lz4,atime=off,/dev/disk/by-id/only, monthly scrub, snapshots 24 h / 7 d.
bulk is exported over NFS by a guest pinned to pve-03. Reproducible data only.
4. Student homes over NFS
A student's home follows them: the same files whether they log into a lab VM on carnaval, a
bare-metal machine from the pool, or a course VM they were given. That only works if the home lives
in exactly one place and everything else mounts it. Today that place is the FS2500; in the target
design it is filer-01, a machine that does nothing else (decided 2026-08-16, closing the earlier
A/B question). The set is small — 71 homes, 428 GB.
What a student sees does not depend on the option: their Linux account inside a VM has a real
home on the VM's own disk (/home/firstname.lastname, disposable, dies with the VM) and a
symlink ~/nas_home to the NFS mount, which is the thing that survives. The rule to give them is
unchanged: anything you want to keep goes in ~/nas_home.
The filer
filer-01 is the twin of the pbs-01 chassis — the two R740xd of the buy list,
bought together — racked in rack-A, on sw-100g, standalone PVE host outside the cluster.
| Layer | Where it is |
|---|---|
| The machine | Dell R740xd 24 × SFF all-NVMe — same chassis as pbs-01: one spares pool, each the other's cold-spare chassis |
| The data | ZFS homes pool, 2 × PM983 mirror from the surplus ≈ 1.7 TiB (4 × today's set) — grown by adding mirror vdevs, 20 bays free |
| The server | the filer host itself exports over NFSv4.1 — no guest in the path: kernel NFS in an LXC needs a privileged container, and with no cluster there is no HA address a guest would carry |
| The export | filer-01:/export/homes/firstname.lastname, in the storage zone |
| Quotas / snapshots | per-user ZFS quota (mechanism per the §10 item) · snapshots 24 h / 7 d |
| Backup | nightly proxmox-backup-client job from the filer host into pbs-01 (307) · monthly cold export to the FS2500 |
| Inside a lab VM | mounted at /exports/firstname.lastname, with ~/nas_home pointing to it — the convention students already know from Calypso and carnaval |
It takes both halves of the trade the two earlier options split. From "FS2500 stays primary":
the sec=sys blast radius stays a box that holds nothing but homes — no rule from the lab zones
into the cluster that runs GitLab. From "dataset on the production cluster": ZFS end to end —
quotas, snapshots, a PBS chain — and DSM stops being the one non-ZFS storage OS holding a
primary. What it costs against those options: one more machine to know, and no HA — the homes
depend on one box. Accepted: the mount is non-fatal (a VM boots with an empty ~/nas_home), the
data has a nightly copy in the other building, and the box rebuilds onto the pbs-01 chassis if
it dies.
The migration is one rsync preserving numeric ownership from
nas:/volume1/isc3_homes/homes into the new pool, a second pass to catch changes, then the
mount flip in the VM templates — phase 4 (§11), with the FS2500 export
left read-only for a semester as the fallback.
Mechanics that must not drift
- One subdirectory per client, not the whole export. A lab VM mounts only that student's
directory; another student's path does not exist inside the VM. The mount is set up by cloud-init
when the VM is created, and it is non-fatal — a VM whose NFS server is unreachable still boots,
with an empty
~/nas_home. - The numeric UID is the authorisation. NFS compares numbers, not names, so an account created with the wrong UID mounts successfully and then cannot read its own files. The UID register is the source of truth and the provisioning script reads it.
sec=systrusts the client, so root on a lab VM is root over the cohort. Anyone with root in the lab zone can become any UID. That is whysudoon a lab VM is granted by name, and why the export is reachable fromlab-virtualandlab-metalonly through an explicit rule into thestoragezone — never fromsrv-*ormgmt. Kerberos (sec=krb5p) would remove the assumption; it is not planned for the first pass.- Per-user quota. One dataset per student with a
quota, or one dataset withzfs userquota(§10); the rule does not depend on the mechanism — one runaway job must not fill the volume every other home shares. - The 307 side mounts across the fibre. The filer sits in building 19, so the pool and the
benches reach it over the inter-building link. Acceptable for editing and building, not for
datasets — those belong in
bulkor on the node's local disk.
5. Backup
| Leg | Box | Pool | Cadence |
|---|---|---|---|
| Replication | the four production nodes on sw-100g | pvesr, ZFS | 5–15 min — availability, not backup |
| Backup | pbs-01 — Dell R740xd 24 × SFF all-NVMe in 307, delivered | 8 × PM983 1.92 TB RAIDZ2 ≈ 10.5 TiB · boot on the BOSS 2 × 240 GB M.2 mirror · 16 bays free | nightly vzdump of every guest over the fibre + the homes from filer-01 (proxmox-backup-client) · verify weekly · prune 14 d / 8 w / 6 m |
| Cold | nas FS2500 | export of configs and homes | monthly |
One box, not the pull-synced pair of the earlier draft — the second chassis is the
homes filer. What the pull bought and how it is replaced, at zero cost: the token
the PVE nodes hold is DatastoreBackup-only — it can write new snapshots but neither prune
nor delete existing ones — and the datastore sits on a ZFS dataset with its own snapshot
schedule, so a compromised node still cannot destroy history. A dead pbs-01 means no backup
until the chassis is replaced; the filer is the same model, so the cold-spare chassis exists. A
night with the fibre down is a missed run — PBS retries.
Sizing: on 2026-08-14 the PBS datastore holding all 11 rumba
guests under the current retention was 171 GiB deduplicated (srv-pbs, 5 TB store, 4 % used);
the raw vzdump copy on rumba is 1.2 TiB. Adding what this design brings into the estate — the
homes dataset (428 GB today), GitLab, the Calypso services — gives a couple of TiB of
primary data, three to six once 14 d / 8 w / 6 m retention is applied. 10.5 TiB is a
three-year target with room, and the growth path is bays, not a new machine.
Growth, in the order it should be taken:
| Step | What it costs |
|---|---|
| 4 more PM983 (a second RAIDZ2 vdev) | free — from the spares |
| Fill the remaining bays with SATA SSDs | market price, any 2.5" drive fits |
| A Dell SC220 shelf, 24 × 2.5", 6G SAS | €42 net, 13 in stock, plus an external SAS HBA |
| An LFF shelf with 14 TB spinners | no Dell shelf is listed — MD1200/MD1400 are both empty. Scenario 4 specifies the step on an HPE D3600 instead, ≈ €4 000, every line in stock |
pbs-01 sits in the lab room but on the mgmt VLAN, never the lab zone, and is NUT master on the
2 kVA UPS. Restore test once per semester, into a throwaway VM, written down. The homes are the
one thing in the design that cannot be rebuilt from a playbook; they have a copy in both
buildings — primary on filer-01 (19), nightly on pbs-01 (307), monthly on the nas.
6. Network
| Fabric | Equipment | Carries |
|---|---|---|
| WAN + zone firewall | ccr-core — CCR2216 · WAN on SFP28 1 · QSFP28 1 → sw-100g at 100 G | NAT, 80/443 + VPN UDP in, WireGuard break-glass, stateful default-drop between zones — only WAN and inter-zone traffic crosses it. The second CCR2216 sits in the rack unpowered as a cold spare, ccr-core's export in git |
| 100 G fabric | sw-100g — CRS520-4XS-16XQ, ordered 2026-08-30 · 16 × QSFP28 + 4 × SFP28 · L3 HW offload | pve-01/02/03 and gpu-01 at 100 G, filer-01 too — pvesr, live migration, NFS, VM VLANs; intra-trust routing at line rate; 9 QSFP28 ports free |
| 25 G server access | sw-100g SFP28 1 | mgmt-01 alone — ConnectX-4/5, 3 SFP28 ports left |
| 10 G aggregation | sw-10g — S5850-32S2Q, 40 G up to sw-100g | carnaval0…7 (8 of the 32 ports), growth |
| 1 G access | sw-access — S3900 | corosync ring0 for both clusters (ring1 on a 25 G VLAN), misc 1 G |
| OOB | sw-oob — the existing CRS326-24G-2S+RM, 24 × 1 G + 2 × SFP+ | BMC/iDRAC/PDU/UPS only, physically separate |
| 307 | ccr-edge CCR2004 + sw-307 S3600 (4 × SFP+: 2 × nas, 1 × pbs-01, 1 spare) | pool (1 G, PXE), tango, pbs-01, bench |
| Inter-site | ccr-core SFP28 5 ↔ ccr-edge, 10 or 25 G — fibre to qualify, dark or campus | the nightly backups, the homes mount for the 307 side (§4), mgmt VLAN — never corosync |
Cabling
One line = one cable. Port numbers are the plan, not a survey of something that exists.
ccr-core terminates the WAN and is the hub every fabric hangs from — the only box that does
stateful filtering; sw-100g is where the servers
plug in and where the fat internal traffic is routed in hardware without going back up.
No cable appears in the buy list — the DAC stock on hand runs from 10 G to 200 G HDR.
The same thing as a patch list — this is what to take to the rack:
| From | Port | To | Port | Cable | Speed |
|---|---|---|---|---|---|
| SInf drop | — | ccr-core | SFP28 1 | per SInf | 1–10 G |
ccr-core | QSFP28 1 | sw-100g | QSFP28 16 | DAC | 100 G |
ccr-core | SFP28 3 | sw-access | SFP+ 1 | DAC + 10 G module | 10 G |
ccr-core | SFP28 4 | sw-oob | SFP+ 1 | DAC + 10 G module | 10 G |
ccr-core | SFP28 5 | ccr-edge (307) | SFP+ 1 (an SFP28 if 25 G) | inter-building fibre | 10 or 25 G |
sw-100g | QSFP28 1–3 | pve-01, pve-02, pve-03 | ConnectX-6 port 1 | DAC | 3 × 100 G |
sw-100g | QSFP28 4 | gpu-01 | ConnectX 100 G port 1 | DAC | 100 G |
sw-100g | QSFP28 5 | filer-01 | ConnectX 100 G port 1 | DAC | 100 G |
sw-100g | QSFP28 6 | sw-10g | QSFP+ 1 | QSFP28 → QSFP+ | 40 G |
sw-100g | SFP28 1 | mgmt-01 | ConnectX-4/5 port 1 | DAC | 25 G |
sw-10g | SFP+ 1–8 | carnaval0…7 | ConnectX-4/5 port 1 | DAC | 8 × 10 G |
sw-access | 1 G 1–6 | pve-01/02/03, gpu-01, filer-01, mgmt-01 | onboard NIC 1 | Cat6 | 1 G |
sw-access | 1 G 7–14 | carnaval0…7 | onboard NIC 1 | Cat6 through the patch panels | 1 G |
sw-oob | 1 G 1–6 | BMC pve-01/02, iDRAC9 pve-03, BMC gpu-01, iDRAC9 filer-01, iDRAC9 mgmt-01 | dedicated BMC port | Cat6 | 1 G |
sw-oob | 1 G 7–14 | iDRAC8 carnaval0…7 | iDRAC port | Cat6 through the patch panels | 1 G |
sw-oob | 1 G 15–17 | 2 × PDU, UPS network card | — | Cat6 | 1 G |
ccr-edge | SFP+ 2 | sw-307 | SFP+ 1 | DAC | 10 G |
ccr-edge | SFP+ 3–4 | nas | 10 G ports 1–2 | DAC, bonded | 2 × 10 G |
ccr-edge | SFP+ 5 | pbs-01 | ConnectX-4 Lx port 1 | DAC | 10 G |
sw-307 | 1 G ×16 | pool-0…7 + their iDRAC8 | onboard + iDRAC | Cat6, two VLANs | 1 G |
sw-307 | 1 G ×1 | pbs-01 iDRAC9 | iDRAC port | Cat6, oob VLAN | 1 G |
sw-307 | 1 G ×2 | tango0, tango1 | onboard | Cat6 | 1 G |
tango0 | QSFP 1 | tango1 | QSFP 1 | DAC, no switch | 200 G |
Port budget after this: sw-100g has 9 free QSFP28 and 3 free SFP28, sw-10g 24 free
SFP+, ccr-core 8 free SFP28 (SFP28 2 included) and one free QSFP28, ccr-edge 7 free SFP+, sw-307 all 4 SFP+
free. Cables are not a line item: the DAC stock covers 10 G to 200 G HDR. Every rack-A machine,
mgmt-01 included, is recoverable from the OOB island — one reason calypsomaster took the
mgmt-01 role over rumba, which has no BMC.
Zones, default-drop between all of them:
| Zone | Contents | Rule |
|---|---|---|
oob | BMC, iDRAC, PDU, UPS | admin VPN group only |
mgmt | PVE/PBS UIs, corosync, MAAS, pbs-01 | admin VPN + admin workstations |
srv-public | reverse proxy | 80/443 from the WAN interface of ccr-core, nothing else |
srv-internal | GitLab, Keycloak, DBs | via proxy + admin SSH |
lab-virtual | carnaval guests, CI runner VMs | outbound + 443 to GitLab + NFS to the homes export; nothing toward srv-*/mgmt/oob |
lab-metal | pool, bench, tango | untrusted = Internet; outbound only, one path to MAAS, NFS to the homes export |
storage | the homes export on filer-01 (§4), PBS, replication | internal only |
7. Sites, power, racks
| Building 19 | 23N307 | |
|---|---|---|
| Cooling | 22 kW chilled water | 3 kW chilled water |
| Feed | 1 × IEC 60309 32 A 3ph (22 kVA) + 2 × 16 A 3ph (11 kVA each), surveyed on site 27 August 2026 — one 16 A way per rack, the 32 A way for the UPS; the room has its own board, so further ways can be requested (electrical request) | — |
| Load | ~4.2 kW typical (rack-A ~2.7 with gpu-01 idle · rack-B ~1.5 with the 8 lab nodes) · ~6.5 kW with the GPUs and the lab cluster both loaded | ~0.9 kW continuous · ≤ 2.6 kW in session (pool on switched outlets) |
| UPS | Eaton 9PX 8000i 3:1 (7.2 kW, 3-phase in on the 32 A way, network card, bypass), to buy — feeds rack-A and both cooling modules, ordered shutdown tested | existing 2 kVA — pbs-01, nas, edge |
| PDU | 2 × HPE G2 Metered 11 kVA 0U, to buy — load-segment metering, one per rack | AP8681 stays — per-outlet metering enforces the pool budget |
| Alarms | rack inlet/outlet + water flow, Telegram + mail | room alarm at 30 °C |
Rack-A with gpu-01 loaded and both cooling modules sits at ~4.5 kW; the 7.2 kW unit carries it
with margin and holds ~11 min at 5 kW, ~17 min at 3.5 kW on internal batteries. The 3:1 input
takes ≈ 12 A per phase at full load, so it wires onto the existing 32 A three-phase way and no
single-phase way has to be added to the board.
Decide with the measured figures.
Every machine is dual-PSU, so a third PDU would give rack-A an A/B feed and let it survive losing one arrival — but two PDUs cannot share one CEE socket, so it needs a second socket, taken from the room's board on the other rack's way. That is an electrical-request decision, made in phase 0 before energisation, not a retrofit. Rack-B is disposable by design and stays single-fed either way.
The physical installation plan for building 19 — dimensioned floor plan, section with mounting heights, and the works packages per trade, in French for the contractors — is on implantation bâtiment 19.
8. Services
| Service | Where |
|---|---|
| Reverse proxy (only inbound VM), GitLab CE + runners, Keycloak (edu-ID), NetBird VPN, mail relay, DBs | production cluster, HA |
| Student homes over NFSv4.1 — §4 | filer-01, standalone |
| PDM, NOC, SOC, MAAS/PXE, rack status page, corosync QDevice | mgmt-01 |
| Student VMs/CTs, CUDA on the A2/T4 nodes | carnaval |
runner-gpu-0x — gitlab-runner, docker executor, one RTX PRO 4500 passed through per VM, tag gpu | gpu-01 — interim: T4 on carnaval2 · rebuilt from playbook, not backed up |
| Course inference / training | gpu-01, spare GPUs |
| Per-course bare metal (Kubernetes, Ceph, Slurm) | pool, via MAAS |
Moodle is absent by design: ISC Learn moves to managed hosting, off the rack, so the filière's most critical service survives a site loss.
9. Buy list
Prices read from ServerShop24 on 14 August 2026, net, excl.
VAT and shipping; the shop displays CHF and Swiss shipping with
?currency=CHF&ShipToCountry=4 on any product URL, and an export to Switzerland is invoiced
without German VAT (Swiss import VAT and duty come on top). Stock moves — confirm every line
against a live quote.
| Item | Qty | Unit | ≈ total | For |
|---|---|---|---|---|
| Dell R740xd 24 × SFF all-NVMe (RC9) — 2 × Xeon Gold 6244, 32 GB, iDRAC9 ent., 2 × 1100 W — delivered August 2026, RG-789539 | 2, identical | €1 639 | €3 277 | pbs-01 (backup, 307) + filer-01 (homes, 19) |
| PBS + filer disks — bought with the chassis | 6 PM983 + 3 BOSS | €294 · €126 | €2 143 | 8 PM983 in pbs-01 (§5), 2 in filer-01 (§4), the rest from the surplus · boot on BOSS 2 × 240 GB M.2 (one card per server + one spare) · also 36 carriers, 26 blanks, 3 PSUs, 4 fans, ≈ €1 065 |
pve-03 RAM | 0 | — | €0 | 256 GB taken from epyc0, which was populated to 1 TB (log) |
| DAC cables, all speeds | — | — | €0 | stock on hand, 10 G to 200 G HDR |
sw-oob | 0 | — | €0 | the CRS326-24G-2S+RM already in the rack |
| 100 G NIC | 0 | — | €0 | the ≥ 6 ConnectX-6 spares cover pve-03, gpu-01, filer-01 and pve-01 |
| 25 G NIC | 0 | — | €0 | ConnectX-4 Lx stock covers all 10 lab, mgmt and pbs-01 links |
| Subtotal, base | ≈ €6 500 | |||
| Gigabyte G242-Z11 — EPYC 7443 24 c, 32 GB, 4 × PCIe Gen4 x16, 2 × 1600 W (13 in stock, September 2026) | 1 | €1 538 | €1 538 | gpu-01 chassis |
32 GB PC4-3200 RDIMM for gpu-01 | 8 | €294 | €2 353 | 256 GB, one module per channel; the chassis ships 2 × 16 GB, which go to the shelf |
| NVIDIA RTX PRO 4500 Blackwell Server Edition 32 GB — new, 165 W passive, dual-slot (3 in stock, September 2026) | 3 | €3 361 | €10 084 | one per runner VM + inference |
Subtotal, gpu-01 | ≈ €14 000 | |||
| MikroTik CRS520-4XS-16XQ-RM (new, Swiss retail) — ordered 2026-08-30 | 1 | ~CHF 1 700 | sw-100g | |
| HPE G2 Metered 11 kVA 0U — P9R61A / 870304-001, IEC 60309 16 A 3ph, 36× C13 + 6× C19, new (18 in stock) | 2–3 | CHF 215 | CHF 430–645 | building 19 racks; the third gives rack-A an A/B feed |
| UPS — Eaton 9PX 8000i 3:1 RT6U HotSwap Netpack (9PX8KIRTNBP31): 8 kVA / 7.2 kW, 3-phase hardwired in, 6U, network card + bypass included — proposed 2026-08-30 | 1 | CHF 3 991 incl. VAT (Bechtle CH) | rack-A + cooling, on the 32 A three-phase way · the 11 kVA / 10 kW sibling (9PX11KIRTNBP31) is CHF 6 468 |
What the market check changed, against the figures the earlier drafts carried:
- No Dell 12 × LFF exists in the catalogue today — not the R740xd (both LFF categories read 0 items), not the R730xd, not the R7515, and not as an LFF shelf either (the MD1200/MD1400 listings are empty; the only expansion units stocked are 2.5"). The in-stock Dell LFF machine is the R750xs 8 × LFF at €4 135 net. Since the store is 171 GiB and the surplus PM983 cover it sixty times over, the answer is the 24 × SFF R740xd — in stock, Dell, and it takes the drives already on the shelf.
- Take the RC9 listings, not the cheaper RC1. The drives are PM983, which are U.2 NVMe: the RC1 machines at €1 081 carry a SAS/SATA backplane and would need PCIe carriers to see them, while the RC9 all-NVMe variant at €1 639 (17 in stock) has a backplane of 16 × U.2 plus 8 dual-mode SAS/SATA/U.2 bays — every bay usable, the boot mirror on the BOSS M.2 pair, and no HBA in the path, which is what ZFS wants. €558 more for the pair, and no adapter to debug.
- 2.5" spinners are not an alternative. The largest 2.5" SAS listed is 2.4 TB at €252, 105 €/TB against 18 €/TB for a 14 TB LFF at the same price. If bulk retention ever appears, it arrives as an LFF shelf behind the R740xd, not as 24 small disks inside it.
- No spinner is bought now, which also retires the 8 TB line (out of stock, €109) and the 14 TB line (€252, 78 units) — kept in §5 as the growth price.
- DDR4-3200 RDIMMs cost €294 each, not the €50–70 the €400–550 line for eight modules assumed.
pve-03no longer buys any —epyc0turned out to hold 1 TB in 32 slots and gave up sixteen modules — so the line lands once, ongpu-01. Buy only what each machine needs to do its job. - No PM1735 3.2 TB spare is available — the U.3 version is listed at €319 but out of stock, and the only AIC in stock is a Lenovo 6.4 TB at €1 008. The EPYC nodes therefore run without a cold spare until one appears; watch the listing rather than budget for it.
- The G242-Z11 listing to take is the EPYC 7443 one, €1 538, 13 in stock (September 2026).
The page carries four Milan variants — 7443 24 c, 7513 32 c, 73F3 16 c, 7713P 64 c — and they
all expose the same CPU flags, so with guests on
cpu: x86-64-v3compatibility does not discriminate between them. The 7443 is taken for two other reasons: at the same 200 W it measures ~18 % ahead of the 7513 in single-thread for a ~4 % aggregate deficit (PassMark, September 2026), which suits CI jobs whose critical path is sequential; and it is not aPSKU, so one spare chip coversgpu-01and both EPYC nodes (the spare R282-Z92 carries Rome 7302s). No drives are included, but the two rear bays take U.2 NVMe, so the mirror comes from the PM983 spares. The chassis normally comes with 8-pin GPU cables; the seller confirmed it will be supplied with 12VHPWR ones instead (September 2026), so nothing extra has to be sourced to fit the cards. - The RTX PRO 4500 Server Edition is dual-slot, not single-slot — 165 W, passive, one 12VHPWR connector. Four still fit the chassis, but they fill it.
- The PDU line drops from CHF 1 500–2 500 to ~CHF 215 apiece, new and in stock. The listing states IEC 60309 32 A on both the 11 kVA and the 22 kVA article, and only the 22 kVA (P9R87A) is 32 A: 11 kVA at 400 V is 15.9 A per phase, and HPE lists the P9R61A as 60309 3P+N+E 16 A — which is what building 19 provides. Get the fitted plug confirmed in writing before ordering.
Offsets: ~20 surplus PM983 (≈ CHF 1 800–2 600) and the R630s left over after five
join carnaval. calypsomaster left the sale list on 2026-08-20 — it becomes mgmt-01.
10. Open items
| Item | Blocks |
|---|---|
filer-01 homes layout — one dataset per student with a quota, or one dataset with zfs userquota (§4) | the quota mechanism and the export tree, before the phase-4 migration |
A DatastoreBackup-only PBS token — verify it can neither prune nor delete existing snapshots | the one-box push hardening (§5) |
| PM1735 3.2 TB — nothing on the market (the U.3 is out of stock, the only AIC listed is a 6.4 TB Lenovo at €1 008), so both EPYC pools run without a cold spare | watch the listing and buy one when it appears |
epyc0 — ConnectX-6 fitted like its twin? | its sw-100g port (else a card from stock, else €283) |
| Backup growth review each semester — store size against the 10.5 TiB box | when to add the second vdev, then a shelf |
Whole-T4 passthrough on carnaval2 — one passthrough mode per node, VM not LXC | the interim runner |
The 5 spares joining carnaval — PERC batteries, second PSUs, NVMe card per node, and are they the same E5 generation? | the lab cluster's real capacity and its power draw |
| 40 G QSFP+ DAC in a CRS520 QSFP28 port — verify | the S5850 uplink |
| Which flows the CRS520 L3 offload can route vs what must stay on the CCR2216 (stateful rules) | the VLAN/routing plan |
x86-64-v3 guest boots on Zen 2 and Zen 3 — test | fallback x86-64-v2-AES |
| Calypso service wind-down date | frees the 8 pool nodes — calypsomaster itself is already free (inventory) |
rumba's exact lab/provisioning use once its guests migrate | what the chassis and its 8 × 500 GB pool end up doing |
| Measured R630 draw, idle + loaded, at the AP8681 | pool concurrency (4 → 6–8) and the rack-B load line |
| R630 count — 18 racked vs 15 invoiced | how many spares are left once carnaval takes five |
| Inter-building fibre — exists? dark? 10 or 25 G? | PBS sync and the cross-building homes mount; the ccr-core ↔ ccr-edge link is in the design either way |
epyc1 dead channel — CMOS clear, then CPU swap | 480 → 512 GB |
11. Plan d'exécution
En français : c'est le plan de travail partagé avec l'équipe. Le document de séance pour les corps de métier est l'implantation bâtiment 19.
| Phase | Contenu |
|---|---|
| 0 | Commandes — faites pour les deux R740xd (août 2026) ; restent gpu-01 + GPUs, la RAM, sw-100g, l'UPS, les PDU, 4 ConnectX-4 Lx · calypsomaster → mgmt-01 : sauvetage des données résiduelles, H730 en mode HBA, réinstallation PVE par l'iDRAC (plan de conversion) — le QDevice est prêt avant que le cluster existe · le sprint d'inspection, à deux dans le rack : câblage du backplane R7515, variante des ConnectX-6, la carte d'epyc0, le compte des R630 et leur consommation mesurée · dépôt des demandes électricité + fibre pour le bâtiment 19 |
| 1 | Le cœur de production dans le rack actuel, sur les switches existants : pve-01/pve-02 (4 × PM1735 chacun), test x86-64-v3, QDevice sur mgmt-01 · pve-03 rejoint une fois son câblage de backplane vérifié · sw-100g câblé une fois le plan de routage arrêté — le cœur ne l'attend jamais |
| 2 | Migration des invités de service de rumba vers le cluster · rumba vidé part en rôle lab/provisioning, usage exact à définir · GitLab + runners CPU sur le cluster, runner-gpu-01 sur carnaval2 (T4 en passthrough) · pbs-01 racké en 307 (maître NUT) → sauvegardes nocturnes + vérification hebdomadaire |
| 3 | Les racks arrivent (30 octobre) → fit-out terminé, circuit d'eau testé à vide, puis déménagement des contenus rack-A et rack-B au bâtiment 19 · gpu-01 mis en service à la livraison, rejoint le cluster comme quatrième nœud — les runners quittent le T4 · les 5 R630 de réserve rackés en carnaval3…7 |
| 4 | filer-01 racké en rack-A → homes copiés depuis le FS2500 (§4 : rsync en préservant les UID numériques, second passage, bascule du montage dans les templates), export en lecture seule pendant un semestre, FS2500 en tier froid · Calypso s'arrête → 8 × R630 vers le pool avec MAAS + prises commandées |
| 5 | Croissance : cours Ceph/K8s sur le pool, cadence des tests de restauration en place, ports QSFP28 libres remplis au besoin |
Contraintes dures : tester l'étanchéité du circuit d'eau avant de racker quoi que ce soit ; ne pas déménager la production avant que l'alimentation du bâtiment 19 soit sous tension et mesurée.