ISC³ — Target architecture
Author: Pierre-André Mudry · Date: August 2026 · Status: not retained —
scenario 3 was retained on 2026-08-20. §10–13 — the
electrical request, the fit-out lots and the UPS sizing — stay the common reference for every
scenario and are unaffected by the decision.
Scope: the ISC³ teaching datacenter only. epyc0, epyc1 and tango are ISC³ assets. CALC@HEI and the Mellanox MQM8700 are out of scope.
1. What ISC³ will provide
ISC³ is the ISC programme's own datacenter: the services the programme depends on, the machines students are taught on, and the management layer that keeps both observable. It spans two sites — a machine room in building 19 and the lab room 23N307 — under one base layer.
One base layer across the whole estate: Proxmox VE. Hardware is decoupled from purpose. Repurposing a machine means destroying and redeploying guests, never reinstalling a bare-metal OS.
Permanently drawable from the chilled-water loop at building 19 — against ~6 kW of design load
Chilled water in 23N307 — the lab's budget. ~1 kW continuous, ~2.6 kW in session
Building 19 for production, 23N307 for the physical lab and the second backup copy
Usable per production node — two mirror vdevs of 3.2 TB NVMe, against ~4 TB of real need
Saved by running the R630 fleet at ~20 % duty instead of 24/7
1.1 Four workloads
| Workload | What it carries | Requirement | Site | Network zone |
|---|---|---|---|---|
| Services | GitLab + CI, Moodle, public and private web, Keycloak, databases | Stability, HA, backups | 19 — rack-A | srv-public / srv-internal |
| Virtual lab | carnaval student VMs and containers, gpu-01 for CUDA and inference | Disposable, quota'd, rebuilt from templates | 19 — rack-B | lab-virtual |
| Physical lab | Bare-metal pool, multi-vendor network bench, the tango AI pair — and the room visitors are shown | Hands-on, wiped after each course, treated as untrusted | 307 — rack-C | lab-metal |
| Management & monitoring | Proxmox / PBS / PDM, NOC (metrics), SOC (security), MAAS, out-of-band BMC network | Isolation, oversight, has to work when the rest does not | Both | mgmt / oob / storage |
The site split is §3, the network zones §6.5.
Services and student workloads are separated physically, not only by VLAN: different racks at building 19, and a different building for the physical lab. Students cannot, even accidentally, starve or compromise Moodle or GitLab.
1.2 What it delivers
| For | What they get |
|---|---|
| Students | GitLab accounts with CI runners · VMs and containers on carnaval · GPUs for CUDA and inference · eight bare-metal nodes they provision themselves (Kubernetes, Ceph, Slurm) · a multi-vendor network bench · home directories with snapshots and a copy in another building |
| Courses | Labs provisioned per course over PXE/MAAS and wiped afterwards · guests built from templates rather than by hand · Moodle |
| The programme | GitLab CE, Moodle, public and private web hosting behind a single reverse proxy · SWITCH edu-ID sign-on through Keycloak · self-service VPN with access by identity group |
| Operators | One Proxmox pane over both clusters · 3-2-1 backups with the second copy in the other building · metrics and alerting (NOC), security monitoring (SOC) · out-of-band access to every BMC |
| Visitors | 23N307 — a lab room with students working in it, plus a wall dashboard of the estate |
1.3 The two sites, and the numbers that size them
- Building 19 — the datacenter. Two 42 U water-cooled racks on order: EPV
LECR13-MW-G-8-12-42A, 800 + 185 × 1200 × 2100 mm, 13 kW of water cooling each, acoustically damped, max equipment depth 950 mm (confirmation 2026-40028, delivery 30 October 2026). 22 kW is drawable permanently from the chilled-water network, against ~6 kW of design load. The room is bare apart from the cooling — the electrical request is §10, the fit-out §11. - 23N307 — the lab. The existing 42 U rack, 3 kW of chilled water, in an occupied room students can walk into. 3 kW is the number that sizes the lab: the rack measured 1.1 kW typical, 3.6 kW peak in July and the room reached 76.4 °C on 23 June — the 3.6 kW peak exceeded the room's 3 kW of cooling.
Both cooling figures are hydraulic capacity, not air conditioning, so the CoolRacks can be plumbed. Two buildings also give ISC³ a second failure domain, which is why the second backup copy lives in 307.
2. Six structural decisions
The second copy is in the other building
ISC³ has never had a backup outside the room it protects. pbs-02 in 307, synced by pull from building 19: different building, different power, different cooling. A fire or a flood at 19 no longer takes the backups with it. Cost: one repurposed server.
Physical lab in 307, sized by its 3 kW
The bare-metal pool, the network bench and the AI pair sit in 23N307, because touching them is the objective. carnaval and gpu-01 stay at 19: 1.4 and 2.5 kW against a 3 kW room, and nothing about a VM reached over a VPN gains from being down the hall.
The R630 fleet: a managed duty cycle
Every machine works, and node count teaches what three fat nodes cannot. Eight R630 running 24/7 cost ≈ CHF 5 900/year; the same eight at 20 % duty cost ≈ 1 200. Hence eight in a bare-metal pool, outside the cluster, off by default — with N powered at once set by measurement at the PDU, not by an estimate.
ZFS + replication, not Ceph
A two-person team whose knowledge is not yet shared, only two symmetric nodes, and a 5-minute RPO nobody will notice in a teaching datacenter. Ceph happens on the bare-metal pool, provisioned per course and wiped after — and that experience is what makes revisiting the decision possible.
The cluster does not span the two buildings
Corosync is latency-sensitive, and the inter-building link is not qualified: a quorum depending on it risks fencing on a latency spike. All three nodes stay at 19; the second site carries backups, which tolerate a slow and occasionally absent link.
The ConnectX-6 replaces the 25 G NICs
The fitted single-port card in Ethernet mode at 100 G into the CCR2216's two QSFP28 ports, a second card from the spares as a direct DAC to the other node. The core's two 100 G ports are used exactly once, the six ConnectX-4 Lx already on the shelf stay free for the lab nodes, and the CCR2004 is recycled as ccr-edge in 307.
3. Two rooms, two jobs
| Building 19 — the datacenter | 23N307 — the lab | |
|---|---|---|
| Racks | 2 × CoolRack 42 U — rack-A PROD, rack-B GPU & growth | 1 × existing 42 U rack — rack-C |
| Cooling | 22 kW chilled water, permanent · 26 kW of rack nameplate | 3 kW chilled water |
| Acoustics | Closed, damped racks — noise is not a constraint | Occupied room — noise is a constraint |
| Carries | Production, primary backup, GPU, the persistent lab cluster, the network core | The bare-metal pool, the network bench, the AI pair, the second backup copy |
| Access | Restricted machine room | Where students and visitors go |
| Failure domain | Primary | Secondary — the reason backups live here |
The two buildings are two failure domains at no extra cost: a fire, a flood or a chiller failure in building 19 does not take the backups with it.
3.1 The lab goes to 23N307 — and 3 kW says how much of it
The lab belongs in the room students can walk into. Racking, cabling, PXE booting, watching a node come up, pulling a disk to see what happens — none of that works over a VPN, and all of it is what a programme with Réseaux et systèmes and Sécurité informatique is supposed to teach.
The constraint is that 3 kW does not hold all of it. The split is by whether being physically present matters:
| Goes to 23N307 | Stays at building 19 | |
|---|---|---|
| Bare-metal pool — 8 × R630, PXE/MAAS per lab | ✅ touching them is the point | |
| Multi-vendor network bench | ✅ | |
tango0 / tango1 — AI pair, quiet, visual | ✅ | |
pbs-02 — second backup copy | ✅ has to be in the other building | |
carnaval0-3 — persistent lab PVE cluster | ✅ student VMs over the VPN; 1.4 kW continuous is 307's entire budget | |
gpu-01 — 4 × PCIe Gen4 x16 | ✅ up to 2.5 kW and a fan wall; needs the 22 kW and the acoustic rack |
Physical labs in 307, virtual labs in 19 — the second column is short because 307's budget is spent on the first. The room also serves as the showroom: a lab room with students working in it.
4. Compute
4.1 Production cluster — three nodes, all in building 19
| Node | Machine | CPU | RAM | Role |
|---|---|---|---|---|
pve-01 | Gigabyte R282-Z92 (epyc0) | 2 × EPYC 7443 — 48 c / 96 t | 512 GB (16 × 32 GB) | HA member |
pve-02 | Gigabyte R282-Z92 (epyc1) | 2 × EPYC 7443 — 48 c / 96 t | 512 GB (16 × 32 GB) | HA member |
pve-03 | Dell R7920 (rumba) | 2 × Xeon Gold 6140 — 72 t | 187 GB | Quorum + Intel-only guests |
Three nodes, three votes, no QDevice. pve-01 and pve-02 are the only live-migration pair (AMD ↔ Intel migration is impossible); the HA group is restricted to those two, and guests that may fail over use cpu: x86-64-v2-AES, never host.
Corosync is latency-sensitive, and the link is not measured: a latency spike would fence nodes. The second site carries backups, which tolerate a slow and occasionally absent link.
Two things to do to the EPYC nodes before they go into production:
- Finish the memory population. All 16 channels per node should carry 32 GB, as one 32 GB module or a pair of 16 GB ones — memory bandwidth on an EPYC 7443 scales with populated channels, on the machines that will carry every production VM.
epyc1is at 480 of 512 GB because one channel does not enumerate; the layout and the defect are on the Epyc page. - Resolve the dead memory channel —
Pon socket P1 ofepyc1, both slots. The earlier note here namingDIMM_P0_B1andDIMM_P1_N1as suspect was wrong: both enumerate and report healthy (August 2026).
4.2 carnaval0-3 — the persistent lab cluster, in building 19
Four Dell R630, one GPU each (2 × A2, 2 × T4), node-local ZFS on the NVMe add-in card, PVE on the PERC array (the BIOS cannot boot the add-in card — verified on 2.17.0 and 2.19.0, not fixable). Everything is built from templates and meant to be destroyed. No HA, no shared storage, no backup — deliberately.
They stay at building 19 rather than moving to the lab room: 1.4 kW continuous is 23N307's entire steady-state budget, and nothing about a student VM reached over the VPN benefits from the machine being down the corridor. They are the virtual half of the lab.
Two hardware fixes before they carry courses: fit the missing PERC H730P batteries (five kits on hand since August 2026) and populate PSU slot 1 on every node — today all of them run on PS2 alone.
4.3 The R630 fleet — duty cycle, not headcount
Every machine in the fleet is in working order. The question is therefore not whether to own them but whether to leave them running.
As general-purpose capacity they are poor value. On throughput one EPYC node is worth roughly 2.5–3 R630 (Zen 3 against Broadwell, plus clock), so the two R282-Z92 already replace five or six of them.
As teaching material the number is different. Node count teaches what three fat nodes cannot: quorum, failure domains, rebalancing, split-brain, PXE, bare-metal provisioning. A ten-node Kubernetes or Ceph lab needs ten nodes. Their marginal cost is zero — they are already here, an E5 v4 R630 resells for a few hundred francs, and a fleet of identical machines is its own parts depot.
The cost is the duty cycle. At ~350 W under load, one R630 is ~3 000 kWh/year ≈ CHF 600 of electricity, ~CHF 730 with cooling.
| Eight R630 | Energy | Cost / year (CHF 0.20/kWh + ~20 % cooling) |
|---|---|---|
| Powered 24/7 | ~24 500 kWh | ≈ CHF 5 900 |
| Powered ~20 % — lab weeks only | ~4 900 kWh | ≈ CHF 1 200 |
The duty cycle is what has to be managed.
| Group | Count | Where | Regime |
|---|---|---|---|
carnaval0-3 — persistent lab PVE cluster | 4 | building 19, rack-B | Always on. Quorum, templates, long-running student VMs |
pool-0 … pool-7 — bare-metal pool | 8 | 23N307, rack-C | Off by default. Switched outlets, PXE/MAAS per lab, wiped after |
| Cold spares | remainder | building 19, rack-B | Unpowered. Parts, PSUs, PERC batteries |
The pool is not a Proxmox cluster member. Nodes that come and go break corosync quorum arithmetic — so the persistent cluster stays at four always-on nodes and the pool sits beside it as bare metal. It is also the better teaching object: a course that installs Kubernetes from scratch, builds a Ceph cluster or runs Slurm on real hardware wants machines it provisions itself.
All eight are racked in 23N307; the room's 3 kW decides how many are powered at the same time. The 350 W figure above is a loaded R630; an idle one in a teaching cluster is likely 150–200 W, which would make six or eight simultaneous nodes comfortable. The AP8681 in that rack meters per outlet — measure a real R630 at idle and under a lab workload before fixing the number. Provisional rule until measured: four at a time; likely six to eight afterwards.
Prerequisites, none of them large: the missing PERC batteries; PSU slot 1 populated; a PXE/MAAS service, which calypsomaster already ran; and Redfish power control from Ansible over the iDRAC8s, which the out-of-band network provides. A lab then starts with one playbook run and ends with another — which is also what enforces the duty cycle.
The count is reconciled (2026-08-17): the invoices account for exactly the eighteen R630 the rack page lists — see the acquisition register.
4.4 GPU
gpu-01 (Gigabyte G242-Z11, 4 × PCIe Gen4 x16) is still to order, and the GPUs with it. It goes in rack-B at building 19 and joins the carnaval cluster, not in the lab room: up to 2.5 kW fully populated is most of 23N307's cooling on its own, and a four-GPU chassis with five redundant chassis fans is not something to sit next to.
Until it exists the lab GPUs are the A2s and T4s in carnaval — adequate for teaching CUDA and inference, inadequate for training, and worth saying so to students. Pick one passthrough mode per node (whole-VM passthrough or shared LXC), never both.
4.5 tango0 / tango1
The two Lenovo ThinkStation PGX go to 23N307, on a shelf in rack-C: quiet, small, ~200 W each, and they demonstrate something people want to see. Their 200 G back-to-back link (measured ~118 Gb/s) comes with them. Reachable over SSH/Jupyter; they are not Proxmox nodes and should not become any.
5. Storage
5.1 The decision: ZFS + replication, not Ceph
The hardware argues for Ceph: 48 NVMe drives, a 200 G link, and Proxmox ships Ceph in the box. Do not build it in production. Three reasons, in order of weight:
- A small team. 213 of 230 commits in this repository are from one person; a system engineer joins end August 2026, which makes it two, with the estate not yet shared knowledge. Ceph's failure modes — a stuck PG, a full OSD, a mon quorum split during a maintenance window — need an operator who has seen them before. ZFS +
pvesrfails in ways one person can reason about. - Two symmetric nodes, not three.
pve-03has one NVMe against 24. A three-node Ceph with one asymmetric member loses redundancy every time you patch a node. - The requirement is soft. A 5-minute RPO on GitLab and Moodle is acceptable; the risk to avoid is an unrecoverable cluster.
Revisit when three symmetric NVMe nodes exist and two people can operate it. The second condition is on its way — a system engineer joins end August 2026 — but the first is not, and shared operating experience takes months to build. Until then Ceph belongs on the bare-metal pool of §4.3 — eight identical R630s in the lab room, provisioned per course and wiped afterwards, is a better place to learn it than a production cluster.
5.2 Pools, and which disks build them
The 19 TB tank on the EPYC bays is dropped. It came from the size of the disk stock, not from a
requirement, and the defective U.2 backplanes turned it from a plan
into a repair project. Added up, what has to fit on the production nodes is:
| Workload | Size |
|---|---|
| Student homes (§5.3) | 428 GB |
| Moodle course backups | 134 GB |
| GitLab + CI artifacts, two years of growth | ~1–2 TB |
| Service disks — Keycloak, web, databases, proxy, NOC, SOC | ~1 TB |
| ISOs and templates | ~0.2 TB |
| Total | ~4 TB |
Two mirror vdevs of 3.2 TB NVMe give ≈ 5.8 TiB per node, so the ~4 TB fits without the bays. That makes the backplane defect an inventory problem rather than a production blocker, and it settles the slot budget the other way: with the bay feeder cards out, the R282-Z92's two x16 and four x8 FHHL slots are all free, which is where the add-in-card NVMe and the ConnectX-6 go.
The two kinds of NVMe on hand then have one obvious machine each, and the interface generations match without waste:
| Stock | Interface | Goes to | Why |
|---|---|---|---|
| 8 × Samsung PM1735, 3.2 TB | HHHL add-in card, PCIe 4.0 x8, 3 DWPD | 4 per EPYC node | The only Gen4 slots in the estate; measured negotiating 16 GT/s x8 on epyc1 |
| 48 × Samsung PM983, 1.92 TB | U.2, PCIe 3.0 x4, read-intensive | 12 into pbs-01 | Its 12 universal bays are Gen3, matching the drives |
| 16 × HPE SAS, 1.8 TB 10K | 2.5" SFF, 12G, 512e | 12 into pbs-01 | The backup workload is sequential and latency-tolerant |
| 2 × SAS 300 GB 10K, shipped with the R7515 | 2.5" SFF, 12G | pbs-01's rpool, if UEFI boot from a universal bay does not work out | Costs two SAS bays instead of the two PM983 in bays 12–13 |
| Site | Machine | Pool | Composition | Usable |
|---|---|---|---|---|
| 19 | pve-01 / pve-02 | rpool | 2 × SATA SSD 480 GB, mirror | — |
| 19 | nvme_pool | 4 × PM1735 3.2 TB, 2 mirror vdevs | ≈ 5.8 TiB | |
| 19 | pve-03 | rpool / hdd | existing mirror / 8 × 500 GB RAIDZ2 | ≈ 2.5 TB cold |
| 19 | pbs-01 — R7515 | rpool | 2 × PM983, mirror — bays 12–13 | 1.7 TiB |
| 19 | backup | 12 × SAS 1.8 TB, 2 × RAIDZ2(6) striped · mirrored special vdev on 2 × PM983 | ≈ 10.5 TiB | |
| 19 | fast | 8 × PM983, 4 mirror vdevs — bays 16–23 | ≈ 7 TiB | |
| 19 | gpu-01 | rpool + scratch | 2 × NVMe, mirror (rear SFF) | ≈ 3.5 TB |
| 19 | carnaval* | fast | node-local NVMe, single device | 0.9 / 0.9 / 0.5 TB |
| 307 | pbs-02 — chassis undecided, see storage option 2 | backup | second copy — pull sync from pbs-01 | sized to match |
| 307 | nas FS2500 | — | cold tier / archive | 1.7 + 6.7 TB |
Rules everywhere: mirrors, not RAIDZ, for VM workloads; compression=lz4; atime=off; /dev/disk/by-id/ paths only; automatic snapshots (24 hourly / 7 daily); monthly scrub. Both EPYC pools carry the storage ID nvme_pool — a different ID on either node means editing guest configs on every migration.
epyc1 today runs three PM1735 as a 3-wide RAIDZ1. Rebuild it as two mirror vdevs when the fourth disk goes in: the pool holds no data yet, so it costs one command, and four in mirrors is both larger (≈ 5.8 against 5.68 TiB) and free of the RAIDZ padding overhead that forced the 16K volume block size.
4 + 4 leaves no spare, on the two machines that carry all production. The PM1735 is out of production, so buy one or two 3.2 TB Gen4 spares while they are still findable — the PM983 stock cannot stand in: the EPYC U.2 bays are dead, and pbs-01's 12 universal bays are the estate's only working U.2 housing.
Losing it destroys the whole backup pool. Two PM983 in bays 14–15, never one.
pbs-01 may be a primaryIt will have the most free space in the estate, which makes it tempting for the student homes — but §5.3 backs the homes up to this machine. Data and its own backup on one chassis is not a backup. Homes go on nvme_pool — 428 GB against ≈ 5.8 TiB per node.
The 36 surplus PM983
Twelve go into pbs-01. For the rest, in order of value: keep 12 as spares; sell or trade ~24, worth roughly CHF 2 200–3 100 at used prices, which covers pbs-01's memory, the sw-oob switch and part of the UPS; keep a handful for the bare-metal pool if a Ceph course wants NVMe, which needs a PLX-switched PCIe-to-U.2 adapter per drive because the R630's Broadwell platform does not bifurcate.
Run the SFF-8654-to-U.2 cable test on an EPYC node once anyway (~CHF 40). It confirms the backplane diagnosis and keeps the repair path open, but do not build a pool on it — the capacity is no longer needed.
5.3 Student homes — close the hole
428 GB of student homes currently live on the FS2500 as a single copy, with no snapshots and no off-box backup, on Toshiba drives with 64 500 power-on hours (7.4 years) that recorded 62–63 °C during the June event. Wear is fine; age is not.
Target: homes move to a dataset on nvme_pool in building 19, served over NFS, with ZFS snapshots, nightly pbs-01, and a pull copy on pbs-02 in 23N307. The FS2500 stops being a primary and becomes a cold tier in the other building — a better job for a 7-year-old array than holding the only copy of anything. The same move retires the srv-pbs-on-Synology-VMM arrangement, which is an expedient and not a target state.
5.4 Backups — 3-2-1, and now genuinely so
pvesrreplicationpve-01↔pve-02, every 5–15 min, over the 200 G cable. Availability, not a backup: corruption replicates.pbs-01in building 19 — nightly, deduplicated, verified weekly, prune 14 d / 8 w / 6 m.pbs-02in 23N307 — nightly pull sync frompbs-01. Pull, not push, so a compromised source cannot delete the copy. Different building, different power, different cooling. This is the layer that does not exist today.- ZFS snapshots as the free safety net against operator error.
Restore test once per semester, into a throwaway VM, written down when it passes.
6. Network
6.1 Fabrics
| Fabric | Medium | Carries | Equipment | Site |
|---|---|---|---|---|
| Service | 25 / 100 GbE | VM and user traffic, NFS, PBS | ccr-core (CCR2216) | 19 |
| Access + corosync | 1 GbE | corosync ring0, carnaval, gpu-01 | sw-access (FS S3900-48T4S) | 19 |
| Out-of-band | 1 GbE | BMC / iDRAC only | sw-oob (managed 1 G, to buy, ~CHF 300) | 19 |
| Replication | 200 G | ZFS send, live migration | no switch — direct cable | 19 |
| Lab | 1 GbE / 10 G | pool, bench, tango, pbs-02 | ccr-edge (CCR2004 reused) + FS S3600 | 307 |
| Inter-site | ? | PBS pull sync, admin | to qualify — see §13 | — |
6.2 The 100 G trick: no new NICs needed
The CCR2216 has 2 × 100 G QSFP28 and the two R282-Z92 each carry a single-port ConnectX-6 VPI — MT28908/MT4123, Gen4 x16, HDR200-capable, firmware 20.31.23.54, read off epyc1's BMC on 2026-08-15 (the card). Run that port in Ethernet mode at 100 G into the two QSFP28 ports: pve-01 and pve-02 each get 100 GbE to the core, the two 100 G ports are used exactly once, and no ConnectX-4 Lx cards need to be bought or found. pve-03 and pbs-01 take 25 G SFP28, off the six ConnectX-4 Lx on the shelf — 25 G is that card's ceiling. The shelf also holds at least six spare ConnectX-6, so a card is available for those two; what is missing is a port — the CCR2216's only two QSFP28 are already spoken for.
The back-to-back link takes a second card per node. This section used to plan it on "port 2", which does not exist — the fitted cards are single-port. It still costs nothing to buy: the shelf holds at least six spare ConnectX-6 of the same model, so each EPYC node takes a second one and a 2 m HDR DAC runs card-to-card, no switch. That link carries ZFS replication and live migration. No configuration, no firmware, no failure mode other than the cable, and replication bursts stay off the fabric corosync and services share. tango0/tango1 already prove the pattern at ~118 Gb/s. What it does spend is a second Gen4 x16 slot per node — confirm against the slot map before committing to it.
Consequence for the rack layout: the 2 m DAC forces pve-01 and pve-02 to be adjacent.
There is no switched 200 G fabric and none is needed — the requirement is exactly one cable between two machines. Re-open the question when a third node needs the same link.
6.3 Border and core are two devices
ccr-border(19) — WAN on SFP28, NAT, default-drop firewall, WireGuard break-glass peers only.ccr-core(19) — L2/L3 for the zones, no WAN interface, no public exposure.ccr-edge(307) — the CCR2004, reused. More than adequate for the lab room, and recycling it means 23N307 costs nothing in new network hardware.- spare CCR2216 — cold, unpowered, both configurations exported to git.
Three CCR2216 are believed to be on hand; the DeepSquare invoice documents one. The design needs two live plus a spare — if fewer turn up, buy, at ~CHF 2 500 per unit.
Retire the CRS326. Not for the reasons first given here — it runs 7.23.3 since 2026-08-09 and telnet is off. The current grounds are a single 800 MHz core and no L3 hardware offload in the flat-bridge configuration (measured ceiling); the case needs restating against what the CCR2216 design actually needs.
6.4 The WAN uplink has to move — or split
The SInf uplink (172.30.7.2, calypso.hevs.ch) terminates in 23N307, which is about to stop being where production lives. Two options, an SInf request either way:
- Preferred — a drop in building 19 terminating on
ccr-border, plus a small separate drop in 23N307 for the lab room. Each site stands on its own. - Fallback — keep the single uplink in 307 and backhaul it to 19 over the inter-building link, which makes that link a production dependency and a single point of failure. Only acceptable if the link turns out to be dark fibre.
Facts to design around: ~4.3 Gb/s aggregate down, ~1.4 Gb/s up; ICMP to the Internet dropped; SMTP limited to 465-to-Infomaniak, a rule SInf granted on request in Aug 2026 (Email) — 25 and 587 stay dropped, and Telegram remains the primary alert channel; inbound is exactly 80/443 + the VPN UDP port.
6.5 Zones, and the end of the flat L2
| Zone | Contents | Rule |
|---|---|---|
oob | BMC / iDRAC, PDU, UPS card | Physically separate switch. No route in or out except the admin VPN group |
mgmt | Proxmox UIs, both PBS, corosync, MAAS | Admin VPN + admin workstations only |
srv-public | reverse proxy — the only VM with inbound Internet | 80/443 from ccr-border to the proxy, nothing else |
srv-internal | GitLab, Moodle, databases, Keycloak | Via the proxy + admin SSH |
lab-virtual | carnaval guests, gpu-01 | Outbound Internet + 443 to GitLab. No flow initiated toward srv-*, mgmt, oob |
lab-metal | pool-0…7, the network bench, tango | Untrusted. Outbound only. Reaches MAAS/PXE on a single allowed path and nothing else |
storage | NFS, PBS traffic, replication | Internal only |
pbs-02 is in the lab room but must not be in the lab zonelab-metal is a room where students install arbitrary operating systems and plug arbitrary equipment into arbitrary ports — it gets the same trust level as the public Internet. pbs-02 sits physically in that room, so it must be on the mgmt VLAN behind the default-drop rule. Backups inside the untrusted zone would undo the point of having them there.
7. Services
| Layer | Choice | Note |
|---|---|---|
| Hypervisor | Proxmox VE 9.x | Production and carnaval |
| Bare metal | MAAS + PXE for pool-0…7 | New. The lab room's provisioning path |
| Ingress | One reverse proxy (Caddy today, Traefik target) | The only VM with inbound traffic. Never a router port-forward |
| Identity | SWITCH edu-ID → Keycloak broker → service | One AAI registration. --autocreate 0 on PVE is the authorisation model |
| VPN | NetBird, self-service, ACLs by identity group | CCR WireGuard kept as 2–3 break-glass peers |
| Git + CI | GitLab CE Omnibus, runners in a separate VM — see the GitLab design | CI runs arbitrary student code: isolation + ceilings |
| LMS | Moodle leaves hannibal for managed hosting, not for the rack — decided 2026-08-11 (plan) | |
| Backup | pbs-01 (19) + pbs-02 pull (307) | §5.4 |
| NOC | Prometheus + Grafana + Alertmanager + Uptime Kuma | 2 VMs. Must monitor two rooms, including 307's total draw |
| SOC | Wazuh + Loki/Promtail + Suricata on a ccr-core mirror port | 2 VMs |
| Console | Proxmox Datacenter Manager | One pane over both clusters |
| Lab room | Status display + live dashboard on the wall | srv-status already exists — give it a screen in 307 |
Nothing currently alerts on a machine going down, a disk failing, or a filesystem filling up. With a closed water-cooled rack in one building and a thermally tight room in the other, that gap matters more — see §9.3.
8. Rack elevations
8.1 rack-A — PROD, building 19
EPV CoolRack, 42 U, 800 + 185 × 1200 × 2100 mm, max equipment depth 950 mm, 13 kW water-cooled, acoustically damped. 23 U used, 18 U free, ~2.4 kW typical.
8.2 rack-B — GPU, virtual lab and growth, building 19
Same chassis. 13 U used, 21 U free, ~1.5 kW typical. Not on UPS, by choice — everything in it is disposable.
8.3 rack-C — the lab, 23N307
The existing rack. 30 U used, 9 U free, ~1.0 kW typical.
- ≤ 1.2 kW continuous, ≤ 2.6 kW supervised during a session. Everything above the baseline sits on switched outlets that default to off.
- How many pool nodes run at once? Measure it. The 350 W figure is a loaded R630; idle is probably 150–200 W. The AP8681 meters per outlet — provisionally four, likely six to eight once measured.
- A lab ends with
ansible-playbook lab-teardown.yml, which is also what removes the heat load. Enforcement is booking, a time window and the automatic teardown. - Room temperature alarm at 30 °C, not 35 — the room has no margin.
pbs-02is in the room but not in its trust zone — see §6.5.- Decommissioning
calypso0-7on its own brings the room back under its thermal budget.
Blank every unused U in the CoolRacks: in a closed rack, an open U is a recirculation path.
9. Power and cooling
9.1 Budget
Building 19 — 22 kW drawable from the chilled-water network, permanently; capacity on that loop is managed by facilities. Rack nameplate is 26 kW, so both racks still cannot run at nameplate at once — irrelevant at the design load.
| Item | Typical | Peak |
|---|---|---|
rack-A — 2 × R282-Z92, R7920, pbs-01, 4 network devices | ~2.4 kW | ~3.9 kW |
rack-B — carnaval0-3 + sw-lab | ~1.5 kW | ~2.0 kW |
rack-B — gpu-01 fully populated (future) | +1.5 kW | +2.5 kW |
| Cooling modules, 2 racks | ~0.6 kW | ~1.2 kW |
Total with gpu-01 | ~6.0 kW | ~9.6 kW |
| Cooling available | 22 kW |
Thermally there is a 2× margin and no open question. The remaining risk is electrical, and §10 is the document to hand over.
23N307 — 3 kW of chilled water. Two numbers, not one, because the load comes and goes:
| Item | Continuous | During a lab session |
|---|---|---|
ccr-edge + sw-307 | 0.15 kW | 0.15 kW |
pbs-02 | 0.30 kW | 0.30 kW |
nas FS2500 | 0.12 kW | 0.12 kW |
tango0 / tango1 | 0.40 kW | 0.40 kW |
| Network bench | 0.05 kW | 0.30 kW |
pool-0…7 — N nodes powered | 0 | +0.7 kW (4 nodes) … +1.6 kW (8 at idle) |
| Total | ~1.0 kW | ~1.9 – 2.9 kW |
| Cooling available | 3 kW |
9.2 UPS
The BlueWalker VFI 2000 RMG PF1 stays in 23N307, where 2 kVA against ~1 kW is right. It still has no network interface and is configured powervalue 0 / MINSUPPLIES 0 — meaning nothing is ever shut down. Fix that, and give pbs-02 the USB link and the NUT master role. The pool is not protected, by design.
Building 19 needs a new rack UPS, ≥ 6 kVA, with a network management card, protecting rack-A and both cooling modules. Graceful shutdown must be configured and tested — otherwise the UPS only delays an unclean shutdown. One controlled discharge test, documented.
The load it has to hold is rack-A (~2.4 kW typical, ~3.9 kW peak) plus both cooling modules (~0.6 / 1.2 kW) — ~3 kW typical, ~5.1 kW worst case. Against 6 kW that is 50 % typical but 85 % at simultaneous peak, so specify unity power factor (6000 VA = 6000 W): the 5400 W variants, including every 3:1 model, sit at 94 % there.
Four candidates, Swiss retail lowest incl. VAT (August 2026). Runtimes are the datasheet figures at ~3 kW, internal batteries:
| Model | kVA / kW | U | ≈ runtime @ 3 kW | Network card | CHF |
|---|---|---|---|---|---|
| Vertiv Liebert GXT5-6000IRT5UXLN | 6 / 6.0 | 5U | 14.5 min | RDU101 included | 3 370 |
| CyberPower OL6KERTHD | 6 / 6.0 | 2U | 6.2 min | RMCARD205 included | 3 420 |
| Eaton 9PX 6000i RT3U Netpack G2 | 6 / 6.0 | 3U | 8.5 min · 38 min + 1 EBM | Network-M3 included | 4 170 (+ EBM 1 800) |
| APC Smart-UPS SRT 6000 RM (SRT6KRMXLI) | 6 / 6.0 | 4U | 9 min | NMC2 included | 4 760 (+ pack 1 730) |
Recommended: the Eaton 9PX G2 plus one 9PXEBM240 (3U + 3U, ≈ CHF 6 000). Unity PF, card in the box, and the free Intelligent Power Manager covers pve-01…03 + pbs-01 for the ordered shutdown §9.3 layer 3 needs; NUT talks to it over netxml-ups. The Vertiv is the cheaper option — CHF 1 400 less, card included, and its 5U internal battery already gives 14.5 min without an extension module.
The supply half reopened with the 27 August 2026 survey: the room has a 32 A three-phase way on top of the two 16 A ways, and its own distribution board. A 6 kVA single-phase unit draws ~30 A at 230 V — no 16 A leg carries that, but a 32 A single-phase way taken from the room's board does, so the four candidates above stay connectable. A 3:1 model on the 32 A three-phase way is the other option: those are the 5400 W variants, where ~5.1 kW worst case sits at 94 %, and they cost about CHF 2 k more than the single-phase units at the same power. Settle the input type with the electrician before ordering.
9.3 Two rooms, two different thermal failure modes
Building 19 — closed rack, fast failure. A 13 kW closed rack that loses water does not take hours to overheat; it takes minutes, because there is no room air to absorb anything. The four-layer plan in thermal protection becomes mandatory, and only layer 1 exists today:
- Warn — per-rack inlet/outlet sensors, water flow and temperature, Telegram and email.
- Shed — automatic power-off of
rack-B's switched outlets above a threshold. - Self-protect — ordered
pveshutdown with guests stopped cleanly, triggered by temperature, not only by UPS state. - Orchestrate — printed procedure taped inside the rack door: open the doors, kill the load. In a closed rack the doors are part of the emergency plan.
Plus a hard trip: PDU-level cutoff on a thermostat, independent of software.
23N307 — open rack, slow failure, no margin. The room has already been to 76.4 °C (incidents). Here the control is the power budget: the pool defaults to off, the PDU alarms on total draw, and the room temperature alarm sits at 30 °C.
10. The electrical request — hand this over as it stands
Building 19 has cooling and nothing else. Every line has a reason; the ones people get wrong are marked.
Feeds
| # | Specification |
|---|---|
| E1 | Two independent feeds, one per rack — 16 A / 400 V three-phase + N + PE, each on its own dedicated breaker at the distribution board. Both sockets exist (surveyed 27 August 2026); what is left to confirm is that they sit on separate breakers |
| E2 | The two feeds on different boards, or at minimum different breakers and RCDs. The point is A/B power for dual-corded servers; behind one protective device the two racks form a single feed |
| E3 | CEE 16 A 5-pole (red), surface-mounted, one per rack, at the rear of the rack position, 1.2–1.5 m above floor. A second socket on one feed is what would let rack-A run A/B — two PDUs cannot share one socket, and it is not retrofittable once energised |
| E4 | Sizing rationale: the two rack feeds are 22 kVA against a design load of ~6 kW (27 %) and a worst case of ~9.6 kW (44 %). What they do not buy is full failover — one feed alone carries the worst case at 87 %, so losing a feed sheds rack-B rather than absorbing both racks. The room's third socket, 32 A three-phase (22 kVA), is the UPS supply and is not counted in that budget |
Protection
| # | Specification |
|---|---|
| E5 | RCD type B, or type A with high immunity explicitly rated for switch-mode supplies. Server PSUs inject DC residual current: a type AC RCD will nuisance-trip or fail to trip. State this in the order |
| E6 | Discrimination: rack RCDs must not be upstream of one another, and neither rack may take the other down |
| E7 | Local earth bar in the room and equipotential bonding of both rack frames — the CoolRacks are closed metal enclosures with water inside them |
| E8 | If local regulation requires it, an emergency stop cutting both rack feeds, at the door, guarded against accidental operation |
Ancillaries
| # | Specification |
|---|---|
| E9 | The rack cooling modules must be fed from the UPS-protected side. Otherwise a mains blip stops the pump and fans while the UPS keeps the servers running |
| E10 | At least 2 × 16 A / 230 V service sockets on the wall, independent of the rack feeds — maintenance, crash cart, vacuum |
| E11 | Working light over both aisles plus emergency lighting, so rear-of-rack work stays possible during an outage |
| E12 | Sub-metering per feed if the board allows it — per-room energy reporting is otherwise guesswork |
| E13 | The rack-A UPS goes on the room's 32 A three-phase socket, hardwired — the candidates in §9.2 all terminate on a terminal block, not a plug. Either a 3:1 unit on the three-phase way (~9 A per phase), or a single-phase unit on a 32 A single-phase way taken from the room's board (~30 A at 230 V for 6 kVA). Give the electrician the model before the board is closed |
Surveyed on site 27 August 2026: 1 × 32 A and 2 × 16 A three-phase sockets, plus the room's own distribution board. The two 16 A ways carry the racks, the 32 A way the UPS, and the design stays within that. The extra ways this section asks for (E3, E10, a single-phase UPS way) are board work, so ask for them while it is open — spare capacity and free breaker positions are worth asking about at the same time.
11. Building 19 — fit-out works list
The room is bare except for the cooling. Lead times, not budget, are the risk — the racks arrive 30 October. Lots 1, 3 and 4 must be requested now; lot 2 must be commissioned and leak-tested before any equipment is racked.
11.1 Lot 1 — Building (request now)
| Item | Why |
|---|---|
| Dust-proof floor finish (epoxy or sealed screed), flatness check | Racks on castors and levelling feet; bare concrete sheds dust into closed cabinets for years |
| Walls and ceiling painted dust-proof | Same reason; to do before the racks are in place |
| Delivery route survey — door width and height, corridor turns, lift capacity and car dimensions | A CoolRack is 2100 × 1200 × 800 mm and ~200 kg empty. Measure the route before 30 October |
| Fire-stopping of every penetration, to the building's compartment rating | Required, and far harder once the cables are pulled |
| Room signage and a rack labelling scheme | The AP8681's outlets still carry factory-default names in 307. Do not repeat that here |
11.2 Lot 2 — Hydraulic (the cooling exists; the connection does not)
| Item | Why |
|---|---|
| Tap-off from the loop, isolation valves per rack, strainer, air vents | You must be able to work on one rack without draining the other |
| Rack hoses and quick couplings, routed so a rack can still roll out | Servicing |
| Insulation of all pipework in the room | Limits condensation on the chilled-water pipes |
| Leak detection — rope sensors under and around both racks, wired to an alarm that reaches a human | Water and 400 V in a closed cabinet |
| Condensate route for the RedBox pump (285 l/h, 4.5 m head, 1 L tank) — connect to the room's existing drain, siphon, fall | The drain exists (surveyed 27 August 2026); its position and height set the head asked of the pump, which is already ordered |
| Flow meter and supply/return temperature probes per rack, readable over the network | Without them, a flow loss or over-temperature water raises no alarm |
| Commissioning and pressure/leak test with the racks empty, signed off | The hard ordering constraint of the project |
11.3 Lot 3 — Electrical (request now)
Items E1–E13, §10. Hand that section over as it stands.
11.4 Lot 4 — Network (request now)
| Item | Why |
|---|---|
| Cable tray from the building's patch room to the rack positions, separated from power runs | The route decides where the patch panels go |
SInf drop terminated in the room, on ccr-border | A request to SInf, not to facilities — and the one with no fallback |
| Inter-building fibre 19 ↔ 23N307 — qualify what exists, pull OS2 pairs if it does not | Decides whether pbs-02 syncs at 10 G or 1 Gb/s |
| Patch panels and optical drawers — OS2 inter-building, OM4 in-room |
11.5 Lot 5 — Safety
| Item | Why |
|---|---|
| Confirm the room is covered by fire detection reported to the building panel; add detectors if not | The racks are closed: an internal fire reaches a ceiling detector late |
| CO₂ extinguisher at the door | |
| Independent room temperature and humidity probe, separate from the rack sensors | So you can tell "the rack is hot" from "the room is hot" |
11.6 Lot 6 — Operations
| Item | Why |
|---|---|
| Clearance — ≥ 1000 mm front and rear of each rack, doors fully opening | The doors are part of the emergency procedure, and a 950 mm-deep server has to come out |
| Workbench, crash cart (screen + keyboard), tool storage | There is no other room to do this in |
| Shelving for spares and packaging | Packaging does not stay on the floor of the room |
| Unpacking / staging area, even temporary, for the 30 October delivery | Two 42 U racks arrive on pallets |
12. What this architecture deliberately does not do
| Not doing | Because |
|---|---|
| Stretch the PVE cluster across the two buildings | Corosync is latency-sensitive and the inter-building link is unqualified (§4.1) |
| Ceph in production | A two-person team, two symmetric nodes, soft RPO (§5.1) |
| Put the bare-metal pool in the Proxmox cluster | Nodes that come and go break corosync quorum arithmetic (§4.3) |
| Leave the R630 fleet powered 24/7 | ≈ CHF 5 900/year for machines idle most of the time. Switched outlets and a duty cycle instead |
Put gpu-01 or carnaval in the lab room | 2.5 kW and 1.4 kW respectively, against a 3 kW room — and neither needs to be touched (§3.1) |
| Keep the CRS326 | A single 800 MHz core routing between subnets in software — the measured ceiling is ~39 MB/s against 110 |
| Buy a new router for 23N307 | The CCR2004 is more than adequate for the lab room |
| Keep the FS2500 as a primary | 7.4-year-old drives holding the only copy of student homes |
Build the 19 TB tank on the EPYC U.2 bays | ~4 TB of real need against ≈ 5.8 TiB from four add-in cards. The bays were surplus, not a requirement (§5.2) |
Put bulk storage or the student homes on pbs-01 | It is the backup target; data and its own backup on one chassis is not a backup (§5.2) |
Put pbs-02 in the lab VLAN | It sits in the untrusted room; it belongs in mgmt (§6.5) |
13. Open items
Nothing here blocks the design; each blocks a purchase, a request or a date. Live action items stay in ops → todo.
To request — long lead times, start now
| # | Item |
|---|---|
| R1 | Electrical at building 19 — the 32 A and 2 × 16 A three-phase sockets exist; what to request is independent breakers, type B RCDs and the extra board ways — §10 |
| R2 | Fit-out works, lots 1 to 6 at building 19 — §11 |
| R3 | SInf uplink terminated at building 19, plus a local drop for 23N307 |
| R4 | Inter-building fibre 19 ↔ 23N307 — qualify, then pull OS2 if needed |
To verify or decide
| Item | Consequence |
|---|---|
pbs-01's 12 NVMe bays — are the backplane-to-board PCIe cables fitted? Ask ServerShop24 for the backplane P/N, or open the lid on arrival | The order confirms the hybrid backplane — 12 SAS/SATA in bays 0–11, 12 universal in 12–23 — but says nothing about the cables, and the whole §5.2 NVMe layout depends on them. The R7515 has only two PCIe slots, so a required extender card competes with the NIC |
pbs-01 memory — 32 GB in two PC4-2666 DIMMs on an 8-channel socket | A quarter of the memory bandwidth, and thin for an ARC over ≈ 10.5 TiB with millions of PBS chunks. 8 × 32 GB DDR4-3200 RDIMM ≈ €2 350: the August 2026 market check found no module under €294, four times what this line used to carry. Pull the 2666 modules rather than mixing them in |
pbs-01 boot device — UEFI boot from a universal NVMe bay | Should work on 14G, but test it at install. The fallback costs nothing — two 300 GB 10K SAS disks shipped with the machine for exactly this — and gives backup as RAIDZ2 6+4 |
| One or two 3.2 TB Gen4 NVMe spares | 4 + 4 leaves the production pair with no spare for an out-of-production part |
| What happens to the 36 surplus PM983 — keep, sell or trade (§5.2) | ~CHF 2 200–3 100 that would cover the pbs-01 memory, sw-oob and part of the UPS |
| Inter-building link — dark fibre, or campus network only? | 10 G or 1 Gb/s for the PBS sync; whether the WAN could ever be backhauled |
| Measured R630 draw, idle and under lab load, from the AP8681 | Sets how many pool nodes run at once — four, or eight. Costs an afternoon |
| CCR2216 count — three believed, one documented | Border/core separation free, or ~CHF 2 500 |
| Supply/return temperature and flow at building 19 (≈ 1.5–2 m³/h per rack) | Undersized flow derates the rack. Capacity itself is settled |
| The two suspect EPYC DIMMs — faulty, or seating? | Handle it during re-population to 16 modules |
| Cold-spare pool: 7 racked spares | |
| UPS size and input type — 6 kVA at 94 % worst case, or 8–11 kVA for ~CHF 1–2 k more? Single-phase on a 32 A board way, or 3:1 on the room's 32 A three-phase socket? | Decides the §9.2 order — upsizing later means replacing the unit, and the input type has to reach the electrician before the board is closed |
Settled: the 22 kW is permanent and outside ISC³'s scope; floor loading and access control at building 19 are not a problem; epyc0, epyc1 and tango belong to ISC³; both cooling figures are chilled water, so the CoolRacks can be plumbed.
14. Order of operations
| Phase | Window | Content |
|---|---|---|
| 0 — Request | now | R1–R4 submitted in writing. Order DIMMs and the OOB switch. Measure an R630. Security-audit remediation. |
| 1 — Prepare in place | Aug–Sep | Build the EPYC pair in the existing rack: PVE cluster, nvme_pool on four PM1735 per node, 200 G back-to-back, pvesr. Migrate rumba's guests onto it. Decommission calypso0-7 — which also brings 23N307 back under its thermal budget for the first time in a year. GitLab CE. |
| 2 — Building 19 fit-out | Sep–Oct | Lots 1 to 6. Electrical energised and tested. Hydraulic connection commissioned and leak-tested with the racks empty, signed off. Delivery route measured. |
| 3 — Racks land and move | 30 Oct → Nov | Install rack-A / rack-B. Migrate in one scheduled window, production last. 23N307 empties of everything that is not lab. |
| 4 — Second site online | Nov–Dec | pbs-01 commissioned — memory ported, the NVMe bays verified and the four pools of §5.2 built. pbs-02 — chassis undecided (§5.2) — commissioned as a pull sync target. sw-oob, VLAN filtering with default-drop, NOC + SOC, new UPS with graceful shutdown and a controlled discharge test. |
| 5 — The lab room | Dec–Jan | pool-0…7 racked and cabled, MAAS/PXE, Redfish power control, switched-outlet enforcement, network bench, console drawer, wall display. First course runs on it. |
| 6 — Grow | 2027 | gpu-01 + GPUs, Moodle off hannibal, Ceph as a course on the pool, off-campus backup copy. |
Commission and leak-test the water before racking anything — a leak during the test must not land on 23 U of production hardware. And do not move production to building 19 before the electrical feed is energised and measured.