The ISC³ datacenter
Date: 2026-08-20 · Status: plan retained, in execution — the two 42 U water-cooled racks are ordered (delivery 30 October 2026) and the two Dell R740xd were delivered (August 2026). The documents: this page is the executive summary — why, what, how. The full design and the execution plan are on Architecture & execution; the alternatives that were weighed are under Considered scenarios; the works-package document for the contractors is implantation bâtiment 19 (French).
1. Why
ISC³ is the ISC programme's own datacenter: the services the programme depends on (GitLab and its CI runners, web hosting, Keycloak identity), the machines students are taught on, and the management layer that keeps both observable. Running it in-house is deliberate: Réseaux et systèmes and Sécurité informatique are taught on the programme's own hardware.
It spans two sites under one base layer, Proxmox VE everywhere: hardware is decoupled from purpose, and repurposing a machine means redeploying guests, never reinstalling a bare-metal OS.
Permanently drawable from the chilled-water loop at building 19 — against ~6.5 kW of design load
Chilled water in 23N307 — the number that sizes the physical lab, enforced at the PDU
Building 19 for production, 23N307 for the physical lab and the backup store — two failure domains
The measured, deduplicated backup store of today's estate — what the backup tier is sized on
Per production node, routed in hardware — one link speed, one card model, cables from stock
2. What
Four workloads, separated physically and by default-drop network zones — students cannot, even accidentally, starve or compromise GitLab:
| Workload | What it carries | Site | Zone |
|---|---|---|---|
| Services | GitLab + CI, public and private web, Keycloak, databases | 19 — rack-A | srv-public / srv-internal |
| Virtual lab | carnaval student VMs and containers · GPU CI runners and course inference on gpu-01 | 19 | lab-virtual |
| Physical lab | Bare-metal pool, multi-vendor network bench, the tango AI pair — the room visitors are shown | 307 | lab-metal |
| Management & monitoring | PDM, NOC, SOC, MAAS, corosync QDevice — on mgmt-01, outside the cluster it watches | 19 | mgmt / oob |
What it delivers: GitLab accounts with CI runners (including GPU) · VMs, containers and CUDA on
carnaval · eight bare-metal nodes students provision themselves (Kubernetes, Ceph, Slurm) · a
network bench · home directories with quotas, snapshots and a nightly copy in the other building ·
SWITCH edu-ID sign-on · one Proxmox pane and out-of-band access to every machine for the
operators. Moodle is absent by design — ISC Learn goes to managed hosting,
so the programme's most critical service survives a site loss.
3. How
The decisions that shape the design — each argued in full on Architecture & execution:
The backup store is in the other building
Production at building 19; the backup store (pbs-01) in 23N307 takes a nightly copy of every guest over the inter-building fibre. Different building, different power, different cooling — a fire or a flood at 19 no longer takes the backups with it.
All-AMD production cluster, four nodes
epyc0, epyc1, epyc3 (the R7515) and gpu-01, guests on x86-64-v3: any CPU-only guest live-migrates to any node. A corosync QDevice on mgmt-01 is the fifth vote, so the cluster survives losing two nodes.
Monitoring outside the cluster
mgmt-01 — calypsomaster, a Dell R740 with an iDRAC9 — runs monitoring, security, MAAS and the QDevice as a standalone host. When the cluster is down, the machine that says so is up, and it is reachable from the out-of-band island.
ZFS + replication, not Ceph
A two-person team whose knowledge is not yet shared, asymmetric nodes, and a 5–15 min RPO that fits a teaching datacenter. A ZFS pool fails one node at a time and says so in one command. Ceph is a course on the bare-metal pool, not production.
The backup tier comes from the disk surplus
The backup store measures 171 GiB deduplicated; the new tier is 8 surplus PM983 in RAIDZ2 ≈ 10.5 TiB — sixty times that, no disk bought. The twin chassis becomes filer-01, the student-homes filer: one spares pool, each the other's cold-spare chassis.
Physical lab in 307, virtual lab at 19
What students must touch — the bare-metal pool, the network bench, tango — goes to 23N307, off by default on switched outlets. What they reach over the network — carnaval, grown to 8 nodes, and gpu-01 — stays at 19, where the cooling is.
4. Where it stands
| Ordered | the two CoolRacks (delivery 30 October 2026) · the CRS520 100 G switch (2026-08-30) |
| Delivered, not yet racked | the two R740xd — pbs-01 + filer-01 (RG-789539, arrived end of August 2026), boot on BOSS M.2 |
| Still to order | gpu-01 + its GPUs, RAM, UPS, PDUs — the buy list, ≈ €14 000–23 000 total depending on the GPU count |
| Requests to submit | building-19 electrical and fit-out, the SInf uplink move, the inter-building fibre |
| The plan | six phases, from the production core in the current rack to the lab room — plan d'exécution (French) |
| Open items | Architecture & execution §10 per blocker · live actions in ops → todo |