Open points / TODOs
Finished items move to the Ops journal so this page only holds live work.
[CRITICAL] / [HIGH] / [MEDIUM] prefix any item with a security or availability
implication (shown as a colored badge), so a cosmetic wish is never mistaken for a data-loss risk.
Unmarked items have neither. Severity is about consequence, not effort — several [CRITICAL]
items are one command. The markers are still plain text under the color: search the page for
[CRITICAL] to get that view.
Security audit — apply the findings
[CRITICAL] — the highest-priority item on this page. A full read-only security audit of the
architecture was run on 2026-08-06. Apply its propositions. The findings, the evidence and the
ordered remediation list live once, in secretzone/security-audit-2026-08.md — kept out of the
published site because it names live weaknesses on named hosts.
Its "Fix today" block is down to one item: the CIFS sudoers rule (F-CODE-1). The fix is written
and waits as PR branch fix/cifs-sudoers-pin (2026-08-28) — the role only reaches the CALC
machines, so the merge is Gregory's call. The rest of the remediation list is still open; several
of its findings already had an item on this page and are marked in place below.
Backups
- Wire the NAS to Healthchecks: the rack-side jobs are wired (2026-08-16, see the service page), the DSM scheduled tasks (scrubs, Hyper Backup if any) are not — needs a curl task in the DSM task scheduler.
- [CRITICAL] Protect the 428 GB of student homes on the
FS2500 —
isc3_homesis still the only copy: no snapshots, no off-box backup. Deferred 2026-08-02 until the Proxmox storage design settles, since rumba'shddpool is the natural target. Two constraints when it restarts:isc3_homesis a real Btrfs subvolume, so DSM 7.2 daily and immutable snapshots are available once Snapshot Replication is installed; and it must be a push — the homes are750owner-only andIscAdmincannot read them, so an rsync pulled from rumba would silently back up nothing. - [HIGH] Commission the new
pbs-01— the chassis question is decided (2026-08-20, retained architecture): one R740xd all-NVMe in 307, delivered end of August 2026, 8 × PM983 RAIDZ2, no pull pair. Next: build the datastore, verify theDatastoreBackup-only token can neither prune nor delete (architecture §5), wire it as NUT master. This is the off-site leg of 3-2-1, so it lands in phase 2, before the rack move. - [HIGH] Tighten the PBS credentials and token scope — see the security audit for what needs moving and rotating.
- [MEDIUM] Schedule a restore drill, quarterly: one CT and one VM from PBS to a scratch VMID, timed and recorded. State an RTO and RPO per service while doing it.
- [MEDIUM] Write the retention and deletion policy (Remi's papers): which data is kept how long, including what happens to a departing student's home. This is also the nLPD gap the audit found on the data-protection side.
- Create backup space for the teachers' laptops (Remi's papers).
- Back up rumba's own services (Remi's papers) — the guests are covered by PBS, the host's service state is not.
Various
-
Confirm the NFS export mode on the device.
nas.mdsaysasyncon the strength of two dated internal notes;sudo cat /etc/exportsneeds a password nobody has entered since. Check on the next NAS session. -
A DNS server off the router? Today every name change means editing the CCR2004 (how it works now). Worth it only if the churn justifies a second thing to run and back up — decide against the alias-domain work already done, which removed much of the churn.
-
Order the Gigabyte G242-Z11 GPU server and its GPUs — the retained plan (2026-08-20) makes it
gpu-01, the fourth production node. Specced 2026-09-01: the EPYC 7443 listing, 3 × RTX PRO 4500, 8 × 32 GB RDIMM — ≈ €14 000 (architecture §9). The card's stock is down to 3 units, so the count is also what is orderable today. The seller confirmed the chassis will come with 12VHPWR cables rather than the usual 8-pin ones. No budget is attached yet. -
Fix the iDRAC state of the existing machines — say which and what. The known cases are the seven R630 spares whose iDRACs were never put on the management LAN (carnaval), and
epyc0's BMC, which needs a reservation and DNS records when it returns. If something else is meant, name it. -
The R7515 —
pve-03, on the bench since 29 August 2026. 256 GB fitted from theepyc0harvest, backplane cabling verified, PVE installed from a USB key (architecture §3). What is left:- Six more PM983 — six are in, giving three mirror vdevs;
nvme_poolreaches its 10.5 TiB with twelve, in bays 12–23. - Create
bulk— 10 × HPE 1.8 TB in bays 2–11, two RAIDZ2(5) vdevs. The other two carryrpool; the 300 GB SAS pair the order sheet promised was not delivered, so the boot mirror costs two of the 1.8 TB and leaves 4 spares instead of 6. - The Mellanox in slot 4 comes up as InfiniBand —
mlxconfig … LINK_TYPE_P1=2to put it in Ethernet mode, as on the other nodes. Model still to identify. - iDRAC housekeeping — the clock reads 1998, which points at a flat CMOS battery: replace it while the lid is off, then set NTP. Address is a bench DHCP lease today and needs a rack reservation.
- Six more PM983 — six are in, giving three mirror vdevs;
-
The PDU answers ICMP but not SNMP, HTTP or HTTPS from a VPN client, while an SNMP read from inside
192.168.88.0/24works — both tested 2026-08-13, from the Mac and fromepyc1. In-rack pollers such as the status page are unaffected, so this only breaks management from a laptop. Decide whether VPN clients should reach PDU management, then either open the path or record it as intended. -
[CRITICAL] NFS configuration issue for students — say which one. This item carries a critical badge and no content. The two candidates already written down are the UID/GID mismatch that leaves five students unable to write their own
~/nas_home, and the mount re-point to192.168.91.250. Either point this at one of them or describe the third thing it means. -
Status-page follow-ups: create a page for upcoming scheduled events (maintenance, outages) linked from the status page, and decide whether a public services-uptime page (status.eduid.ch-style, per service rather than per host) is wanted on top of it.
-
A restricted zone for the ISC admin tooling from git, for two people only — say what "zone" means before building it: a NetBird group and policy, a Proxmox pool, a separate repository, or all three. It touches NetBird and Proxmox either way.
-
Set the OS hostnames of the Tango pair — the docs, the Ansible inventory and the status poller say
tango0/tango1; the racked node reports the factory namespark-4425. -
Give the racked Tango node its documented
192.168.88.20— it sits on a dynamic192.168.88.43because its live MAC4C:BB:47:2F:44:25has no reservation, while.20/.21are reserved for two other MACs. Decide which port is meant to be the LAN one, then either recable or move the reservation; the status poller and Ansible inventory both expect.20. -
Review Change management and Incident management pages, the information is odd — and write the notification templates they refer to. Nothing was lost in the wiki migration: those were red links there too, so the templates have to be written from scratch; the dangling link stubs were removed 2026-08-19.
-
Network follow-ups from the 2026-08-02 audit (applied fixes in the journal, details): 0.
.calypsoDNS domain retired 2026-08-28 (.isc3only, legacy records purged, node FQDNs rewritten). Left: reissue kubeconfigs that namek8s.calypso(the API certificate already listsk8s.isc3); delete the NAS snapshot/volume1/@isc3_homes-pre-renameafter a few days. The guests'/etc/resolv.confwere realigned 2026-09-14 (provisioning/pve/guest-searchdomain.sh, andsrv-pbsby hand); left:srv-test1andsrv-test2, stopped that day, so re-run the script once they are up. VM 107's[special:cloudinit]cache still holdssearchdomain: calypso— its active config saysisc3, so a cloud-init drive regeneration clears it.-
Fast-path route +
nas-fastpath.service— the node half went with the 2026-08-28 rebuild (lab VMs mount.91.250directly). Left: remove the unit from the Ansible role, and decide whether the NAS-siderc.d/nas-fastpath.shroute is still needed now that nothing in192.168.91.0/24talks to.88.250. -
CRS326 — the hygiene pass and the v6 → v7.23.3 upgrade are done (2026-08-04 → 09, what happened). Still open:
- [MEDIUM] L3 hardware offloading is now within reach: the 98DX3236 supports it under
v7, and it is exactly the software-routing ceiling
the 2026-08-02 audit measured (~39 MB/s against 110). Not free — inter-VLAN offload needs a
vlan-filteringbridge with VLAN interfaces, where today all three subnets share one flat bridge, and offloaded traffic bypasses firewall and NAT. - The target architecture still argues for retiring this switch on the grounds of "RouterOS 6.49.19, EOL, telnet enabled" — both of which are now false. The case for replacing it needs restating on current facts.
- [MEDIUM] L3 hardware offloading is now within reach: the 98DX3236 supports it under
v7, and it is exactly the software-routing ceiling
the 2026-08-02 audit measured (~39 MB/s against 110). Not free — inter-VLAN offload needs a
-
-
Pilot the exam VDI proposal (SEB + Guacamole on carnaval) with one class before any real exam — measure guacd load and per-session bandwidth against the sizing figures; the other open questions are listed on the page.
-
Open points specific to the research infrastructure (Chacha, Disco, Mambo, "Dance New") live on the CALC@HEI todo page.
EPYC nodes (epyc0 / epyc1)
Open items for the EPYC pair — epyc1 is racked and running Proxmox,
epyc0 is away in repair. Done work is in the journal and the
history page.
-
[HIGH] Change both EPYC BMC passwords — they are the board serial numbers, printed on the chassis pull-tab and readable by any host root via
dmidecode. Host root therefore escalates to BMC admin (IPMI, KVM, power) on a machine whose IPMI port every VPN user can reach. Record the replacements insecretzone/epyc.md. -
[HIGH] The
epyc1BMC exposes IPMI (623) + Redfish on the flat192.168.88.0/24LAN, reachable by every VPN user — same exposure class as audit finding F-HOST-1. Fold it into whatever the flat-L2 remediation becomes. -
[HIGH] Give
epyc1a working mail path so ZFS and PVE alerts leave the host — postfix has no relayhost and norootalias, so a degradednvme_poolshows only in local mail and the web UI. Smarthost insecretzone/smtp.md(465 to Infomaniak only). -
Storage plan for the EPYC nodes — decided 2026-08-13, nothing executed. Design and the disk allocation across the estate: §5.2. What is left to do:
- Fit the fourth PM1735 in
epyc1and rebuildnvme_poolas two mirror vdevs. The pool holds no data (1.85 MB allocated), so it costs one command — do it before anything lands on the node. Four in mirrors is larger than the current 3-wide RAIDZ1 and has no padding overhead. - Four more into
epyc0when it returns from repair, same layout, same storage ID. - Buy one or two 3.2 TB Gen4 NVMe spares — 4 + 4 leaves the production pair with none, and the PM1735 is out of production.
- Run the SFF-8654-to-U.2 cable test once (~CHF 40, SFF-8654-to-SFF-8639 with an integrated power lead). It confirms the diagnosis and keeps the repair path open; do not build a pool on it.
- Decide what happens to the 36 surplus PM983 — 12 spares, ~24 to sell or trade at roughly CHF 2 200–3 100 (§5.2).
A reseller claim for the backplanes stays open as a separate option; nothing waits on it. The spare R282-Z92's untested backplane is a second one — see the item below.
- Fit the fourth PM1735 in
-
The spare R282-Z92 — 2 × EPYC 7302, no RAM, no drives (config). Record its delivery date and invoice in the register. No CPU to buy: it boots as delivered, at 32 Zen 2 cores against a node's 48 Zen 3, and guests on
cpu: x86-64-v3run on either. What it lacks is memory — the shelf holds 2 × 32 GB, so a swap-in depends on moving the failed node's DIMMs across. Decide whether that is accepted or whether a minimal DIMM set is bought for it. -
Test the spare R282-Z92's U.2 backplane on arrival — it is the only untested 24 × U.2 backplane in the estate, and both production ones are defective. Use the documented LED test (power only, no data cables). A good backplane makes it a fifth option alongside the reseller claim: donate it to
epyc0orepyc1and put the PM983 in bays, or move the drives to this chassis instead. Do this before the chassis is written off as insurance only. -
When
epyc0returns from repair: check the router for a dynamic lease rather than assuming192.168.88.11, give its BMC a reservation andepyc0-bmcDNS records, and verify the pull-tab password insecretzone/epyc.md. Name its ZFS pool with the same storage ID asepyc1's (nvme_pool) or guests will not migrate without config edits. -
Recover or write off the 32 GB missing on
epyc1. Everything short of opening the chassis is eliminated; two physical tests remain, in order — a CMOS/NVRAM clear, then a CPU swap between sockets. Combine them with fitting the fourth PM1735 above. -
Housekeeping on
epyc1:systemctl disable zfs-import@data_pool.serviceretires the unit left behind by the pool deleted 2026-08-12, which fails at every boot and is the node's only failed unit; there are also 60 pending package upgrades. -
While in
epyc1's BIOS setup for anything else, considerMilan0168 Enable AER Cap→Enabledso PCIe link errors become visible (why). Diagnostic value only, and it would not have helped the dead bays, so it is not worth a visit of its own.
Rumba
Open items for the Proxmox node; done work is in the journal and the history page.
- Migrate its guests to the production cluster (phase 2 of the
plan). Rumba does not
join the cluster and does not become
mgmt-01— that role went to calypsomaster on 2026-08-20 (iDRAC9; rumba has no BMC). - Define its lab/provisioning duty once emptied — decided 2026-08-20 that it stays in service
on the lab side; the exact use (bench, provisioning host…) and what its 8 × 500 GB
hddpool carries are open. Its ConnectX-6 returns to the spares shelf. - Give
isc-adm(VM 104,192.168.88.171) an internal DNS name on the CCR2004 (isc-adm/isc-adm.isc3, per the other guests) and decide whether it gets a personal admin account instead of shared root. - Note before any RAM plan: all 24 DIMM slots are populated with 8 GB modules, so memory can only be replaced, never added — details.
- Check
hdd's free space before promising the pool to anything new — several items below want it as a landing spot. - Wire Ansible's dynamic inventory to NetBox (v2 token per
consumer). The inventory itself was seeded 2026-08-16 (
provisioning/netbox/populate.py); still absent from NetBox: cables/connections, the PDU outlet map, device serial numbers.
NAS (FS2500)
Ranked actions from the August 2026 disk-health readout. The 2026-08-02 hardening, first-ever data scrubbing and Volume 2 rebuild are done — history page.
-
Re-point the node mounts to
192.168.91.250. The NAS now has one dedicated 10 GbE port per subnet —eth2=192.168.88.250(DSM,rumba/PBS backups),eth3=192.168.91.250static, in the node subnet (done 2026-08-03; bonding was considered and rejected — redundancy is not the goal, throughput and simplicity are).The node side went with the 2026-08-28 rebuild; what remains is making sure every image and lab-VM template mounts
.91.250, not a per-node script. The rules and both thefstabandpvesm add nfsforms are written up once on the NAS page — how a node must mount it. No export change is needed; the NAS already exports to192.168.91.0/24and matches on the client address.Same edit drops the leftover
syncand the duplicate/exports/line from the nodes'fstab. Neither is urgent: with the exports nowasync, a client-sidesyncwas measured to cost almost nothing (2026-08-05).Afterwards, delete
/usr/local/etc/rc.d/nas-fastpath.sh, the manual192.168.91.0/24 dev eth2route, andnas-fastpath.servicefrom any surviving node. -
Until the nodes are rebuilt, the fast-path route can vanish. Nodes still mounting
192.168.88.250depend on it, and/usr/local/etc/rc.d/nas-fastpath.shre-adds it only at DSM boot, so a link flap costs ~40 % of read throughput silently. Re-checkip routeafter physical work behind the NAS; if the rebuild is far off, anas-fastpath.timer(OnUnitActiveSec=5min) would bound it. -
Tidy
eth0/eth1on the NAS. Both are stillBOOTPROTO=dhcpwith no cable, so plugging either one in re-opens the DSM gateway-election hole described on the NAS page. -
Enable DSM notifications and SSD-lifespan alerts so a degraded array or wearing disk is not discovered by accident.
-
When the 16 GB DIMM arrives (ordered from senetic, still not delivered as of 2026-08-18): install in slot B (→ 24 GB total), then raise the PBS VM to 4 vCPU / 8 GB in VMM and eject its install ISO (Edit → Others — the API can't do it).
UPS and power
The rack UPS has powered rumba and nothing else since
2026-08-03. It is monitored with NUT and shown on the
status page, which alerts on mains loss.
The UPS dropped off rumba's USB bus and there is no remote lever: both ends were reseated, a full reboot reset the xHCI controller, neither produced a kernel event (incident). Until it is fixed a mains failure means rumba runs to a flat battery with no clean shutdown, and the status page shows stale UPS data.
Next time at the rack, in order, watching
ssh root@192.168.88.51 'dmesg -TW | grep --line-buffered -iE "usb 1-|0665"':
- A known-good USB device in port 13 (rear panel, left, upper — where the UPS cable is). It enumerates → the cable or the UPS's own port is at fault; it does not → port 13 is dead, move the UPS to a neighbouring port.
- Swap the USB cable.
- Whatever brings it back, finish with
systemctl restart nut-driver@rackupsandupsc rackups; the driver is not known to reattach on its own.
1. [HIGH] Decide between repairing the USB link and moving to serial
The unit has a DB9 serial port and rumba has two real UARTs, and nutdrv_qx speaks Q* natively
over serial — so the USB cable is not the only possible path, and that link has now failed three
times. The cost is specific: serial carries Q* only, so it loses the
HID interface and with it the genuine RunTimeToEmpty
and RemainingCapacity — exactly the numbers the shutdown decision needs. Decide
the two together.
2. [HIGH] Shut rumba down cleanly before the battery runs out
On a building outage the UPS carries rumba about 55 minutes, then cuts power instantly — a running
hypervisor with live guests and ZFS pools. Nothing prevents that today: upsmon runs
powervalue 0, deliberately. What it lacks is a battery level it can trust, since nutdrv_qx
reports a percentage pinned at 100 % and the correct data sits on a
USB channel NUT cannot read. Three ways to close it:
- A. Trust NUT's
LBflag — one line (powervalue 1,MINSUPPLIES 1). Unverified: nobody knows whether the firmware sets that flag or NUT synthesises it from the broken percentage. If synthesised, the shutdown fires far too early or never. Item 3 settles it. - B. A small shutdown watchdog of our own — a timer reads
ups-read-hidand shuts down under a margin (say 15 min) or on the UPS's genuine shutdown imminent bit. ~50 lines plus a unit, and ours to maintain. Recommended: the only option acting on numbers we have verified. - C. Bridge the good data into NUT via a
dummy-upsdevice kept updated by a script. Stock NUT does the deciding, but it is the most moving parts and needs a staleness guard — a dead updater leavesupsmonreading "battery full" forever.
3. [HIGH] Run one controlled power-fail test
Decided 2026-08-03, not yet run. Runbook: UPS controlled discharge test.
It answers three questions at once — whether option A's LB flag is real, what the battery's
autonomy actually is today (55 min is the UPS's own estimate and cells age), and whether the
shutdown path works end to end. It costs a planned outage; the alternative is learning these
answers during an unplanned one.
4. [MEDIUM] Decide what else the UPS should protect
rumba uses ~14 % of the 2 kVA, so there is headroom, but 2 kVA cannot cover the rack (PDU peak
3.63 kW). The obvious candidate is the NAS, which holds the only copy of the student homes. Needs a
physical cord trace at the rack — and rename the PDU's 24 factory-default outlet names while
there, so the question is answerable remotely next time. Also check whether the chassis has an
intelligent-card slot for a real SNMP interface.
R630 hardware — fit the PERC battery kits
The five kits are on hand (pack with
holder, 070K80 / 0H132V / 037CT1 / 0HD8WG). Confirmed missing (names after the 2026-08-28
renumbering): carnaval1, carnaval2, carnaval7, carnaval8; carnaval5 unconfirmed (it is
also the node with the dead DIMM_B1 — same visit to the room). The per-node readings, the two
hardware traps and how the fleet was swept are on the archived Calypso page.
VPN and identity — NetBird, Keycloak, edu-ID
The remote-access and single-sign-on stack: the NetBird VPN, the Keycloak broker federating to SWITCH edu-ID, and the legacy MikroTik WireGuard being retired behind them. Architecture: §7.
- Legacy WireGuard client confs carry a dead DNS line — the 72 already-issued
.conffiles still sayDNS = 172.30.7.1, which never resolved the internal names (the router-sideclient-dnswas fixed 2026-08-16, history). Fix on next contact — reissue withwg-gen.shor hand-edit toDNS = 192.168.88.1, isc3, calypso; no mass campaign, WireGuard is being retired anyway. - NetBird follow-ups — the deployment, the edu-ID federation (Phase 2) and the admission policy are
done and validated, 2026-08-03 → 06 (journal); the current state lives on the
service page and the
Keycloak page. What remains:
- Onboard users: teachers/admins now; students at the autumn intake, who go straight to NetBird as part of retiring the legacy MikroTik WireGuard.
— done. The claim arrives: as of 2026-09-06eduPersonOrgUnitDNis not being released yetischolds 15 members andhevs16, students among them (*@students.hevs.chaccounts carryRACA-TICO-ISCO, which the 2026-08-05 staff-only measurement could not show). No consumer is left onhes-so: the user gate and the realm's admission gate both testisc.- Verify
replay-config.py --jwt-twinsagainst the post-collapse group model — it pairs<name>/<name>-manualfrom the pre-2026-08-06 identities; the policies now match entitlement jwt groups (vpn-rack-operators,vpn-rack-mgmt) that have no manual twin, so a store rebuild may need it extended before the twins step works. - Optional NetBird group slimming: with a single local account, one break-glass group
suffices — edit
teachers-manual/students-manualout of the four policies first, then delete them (deleting first cascades — incident, which is also why they exist again), keepingadmins-manual;rack-routertags the routing peer but nothing references it — confirm and drop.hes-so/role-rack-adminsstay: the claim carries them, JWT sync would just re-create them. - [HIGH] Require MFA for the Keycloak admin account.
/admin*no longer depends on the client's address — it is behind the admin gate and the grouprole-rack-admins(three members since August) — and the gate authenticates against Keycloak itself. Since 2026-08-05 the same group is also the sole Administrator grant on both Proxmox clusters, so it now carries two systems;root@pamremains the break-glass for that half. MFA on it would be requested per-flow viaacr-values(PVE's realm takes one), which is why REFEDS MFA was deliberately left off at the edu-ID resource level. - Fill
roster.csvfor the autumn cohort — the mechanism is done and proven with one admin, but the file holds a single line. Needs the class list and the address edu-ID actually releases for students; a wrong domain means "logs in fine, no access" for everyone. Verify with the first student login that they also carryRACA-TICO-ISCO(item 2 above). - Re-verify the whole chain with a second identity. Everything so far was proven with one
edu-ID account, which cannot exercise: two people with different roles at once, a user who is
in no role group being refused by
jwt_allow_groups(only reasoned, never observed), and theidp-auto-linkpath for a roster account that has never logged in. A colleague's account or the first student is the test. - Rewrite
roster-sync.shon the Admin REST API (curl+jq) instead ofkcadm.sh. Everykcadmcall starts a JVM (2–3 s), two per rostered person, so a full pass takes minutes; the--only <email>mode added 2026-09-03 brings one arrival down to about 30 s but the yearly move-up and removals still need the full pass. One token request then plain HTTP calls would put the whole file under 15 s. Same behaviour to keep: idempotent, never deletes an account,role-rack-adminsadd-only. - Route
tango0/tango1tovpn-carnaval— they sit in the88subnet, so a host resource per machine innet-carnaval-guestsis the way, and the racked node still runs on a dynamic.43lease (tango): give it its static address first. Decided 2026-09-08: no rule until then, moving the machines physically is an option.
UID allocation — wire the register up
The register exists since 2026-08-05: provisioning/uid/uid-map.csv, seeded from home ownership on
the NAS, issued by uid-alloc.py and made into a home by create-home.sh. Ansible reads it on the
rack; what is left is applying it and deciding about the machines that share the same lists.
- The 13 renumberings are moot — the
/etc/passwdcarrying the drift was wiped with the 2026-08-28 node rebuild. Make the lab-VM templates draw their accounts from the register — that is the item that matters. - The research machines are frozen — Gregory has two PRs open on
dance.yml. Do not touch that inventory oruid_othersuntil they land. Afterwards: theconf/usersid:columns survive only because those lists serve both fleets with different numbering, so deleting them means unifying the two namespaces first — renumber and chown one side. - Point
carnaval-lab-vm.shat the register rather thanstating the NAS. Not urgent: the NAS ownership is what NFS enforces, so it cannot be stale — but it means two code paths for one fact. This one is the code path that stays, so it is the item that matters. - Re-run
uid-alloc.py --verifywhenever the NAS homes change. The nodes that were down when the register was seeded (calypso3–7) were wiped before coming back, so their local accounts never need sweeping. - Simplify
01_users.yml's register logic oncelouis.herederois unified (see the Carnaval item 7 fix). Thecanon/solo/statusbookkeeping inprovisioning/ansible/roles/00_setup/tasks/01_users.yml— one CSV row per number, canonical only when a name has exactly one non-reservedrow — exists solely to represent a name split across two live numbers on two fleets. Once he's down to one number, that becomes the only case it was handling; collapse it back to a plain name → uid dict.
Worth doing before the autumn intake: that is the next batch of numbers to hand out, and the point at which pre-created roster accounts start needing them.
Carnaval (playground cluster)
Ten-node cluster since 2026-08-28 — the cluster page holds the current state. What remains:
-
carnaval5— replace or pullDIMM_B1, then rebuild it (perc-raid1.sh, baked ISO, post-install,pvecm add— the recipe). The missing PERC batteries (carnaval7,8, …) are the fleet-wide sweep in R630 hardware. Also: remove the legacycalypsoNDNS records on the CCR once nothing references them, and register the ten nodes in NetBox. -
Alert when a
fastpool suspends, and decide what happens to the Swissbit cards. Both cards of the pair have now hung the same way —carnaval1on 2026-08-06, unnoticed for 17 h (incident), andcarnaval8on 2026-09-03, unnoticed for five days (incident). The iDRAC cannot see an add-in NVMe, so no SEL entry and no health change. Two separate jobs:- Configure
zfs-zedto mail onzpool suspended/ device removal on every node. Cheap, independent of the planned NOC, and it covers the only failure this cluster has had — twice. - Decide whether the cards stay in service.
carnaval8is powered off holding the one that hung in September (EUI…03550000, 7398 h);gpu9currently runs on the one that hung in August (…05590000, 6555 h, incarnaval9);carnaval7carries a third of the same model. SMART is clean on all of them and was clean before both hangs, so it is no help deciding — the August decision to keep the card rested on exactly that clean report. Check whether Swissbit has firmware newer thanARR50002; there is no warranty case on wear, only on the hang. A card pulled has to be replaced: these are the only 894 GB guest pools in the fleet.
- Configure
-
Kubernetes — phase 2. The shared cluster built 2026-08-28 (bash, not the playbook the design foresaw) has no VMs since 2026-09-18 (
gpu0–2hold the cards;carnaval-k8s.sh uprebuilds it from the current template in ~10 min). Left: edu-ID student access (k3s OIDC → Keycloak, namespace-per-cohort RBAC +ResourceQuota, blocked on item 5), per-team clusters as further instances of the same script, and per-student storage quotas (NAS and local). Student access is wanted for the last block of the Kubernetes course, around late October 2026 (lecturer, 2026-09-09); the fallback is k3s on their laptops. -
Move the guest-facing half to Ansible.
carnaval-lab-vm.shreads who gets in from ISC³-only files since 2026-08-16 (provisioning/keycloak/roster.csv+ssh-keys, decoupled from the Ansible tree calc shares) — and since 2026-08-05 it reads UIDs from the NAS home ownership rather than a login node, so it no longer depends on one at all.community.general.proxmox_kvmdoes what the script does with proper idempotence instead of hand-built cloud-init YAML. The substrate scripts (install, pool, cluster) can stay as they are. The UID columns inconf/users/*.ymlshould be deleted, not repaired, as part of the same job — see UID allocation. -
Grant the cohorts
PVEVMUseron their pools once PVE accounts exist — there are none yet, so there are no student accounts yet. The federation half is done since 2026-08-05 — the cluster takes an edu-ID login through Keycloak, but only for pre-created administrators. What is left is the part that avoids hand-creating ~40 accounts a year:--groups-claim groupsplus--autocreate 1, then the pool ACLs on the claim-derivedstudentsgroup rather than on individuals. -
Then decide which guests deserve a selective backup job — there is deliberately none today. The templates are covered since 2026-08-16 (publish leaves a vzdump on
nas-library); the question remains for infra guests, if the cluster ever grows any worth keeping. -
Five students' UIDs vs GIDs — the Calypso-side half of this is gone with the nodes (2026-08-28); what remains is making sure the lab VMs issue the register's numbers (
provisioning/uid/uid-map.csv) so~/nas_homeis writable — verify on the first cohort VMs. -
Publish the 2026-09-18 CUDA template to
carnaval3,4,6,7,10once they are powered on — they still carry the August image without Slurm/MPI. Per node:qm destroy 910N --purge 1, thenNODES="carnaval9 carnavalN" provisioning/pve/carnaval-guests.sh --cuda-seal(pipeline). -
Collect the 25/26 cohort's SSH keys into
provisioning/keycloak/ssh-keys(format). Every roster person has an account on every lab VM since 2026-08-16, but only 8 of the 34vpn-carnavalhave a key — the other 24 have an account ongpu9with no way into it (2026-09-08). Existing VMs do not pick up roster changes, so the keys have to be in before the VM the cohort will use is built.adrien.reynardnow has a NAS home;loic.christen1needed thenasnamecolumn instead, his home beingloic.christen(fixed 2026-09-08, his account appears at the next provisioning run).
Calypsomaster → mgmt-01
Audited read-only 2026-08-02 and inventoried over SSH and Redfish 2026-08-13: nothing load-bearing
runs on it. Decided 2026-08-20 (retained architecture):
it becomes mgmt-01, the standalone watcher outside the production cluster — PDM, NOC, SOC, MAAS
and the corosync QDevice — taking the role from rumba for its iDRAC9, faster CPUs and free DIMM
slots. The conversion can run now.
Nothing depends on it. Its kubeadm control plane has no user workloads and its only enrolled
workers (calypso9/calypso10) are now carnaval nodes; MAAS 3.6.2 has exactly one machine in its
inventory — itself, DHCP off everywhere; the Docker registry container was started without
published ports, so it has been unreachable all along (5.2 GB of old course images); Zabbix Agent 2
points at a server on epyc0 that no longer exists. Nothing external depends on it either — the
carnaval nodes use router DNS and mount the NAS directly. The only
soft dependency is the NAS-lockout recovery tunnel in the
incident log, and rumba serves that role identically.
What the conversion needs to know. The machine and its disks are specced on the archived Calypso page; the rest is:
- The PERC H730 is hardware RAID — put it in HBA mode before ZFS touches it. U.2 drives do not fit (SAS/SATA backplane), so NVMe has to arrive on add-in cards; the PERC occupies slot 2.
- No recabling: its 10 G DAC lands on the CCR2004
sfp-sfpplus4, samebridge-LANas rumba — already in the admin segment. The idle 1 GbE ports (eno3/eno4) are a free option for a dedicated corosync link later. iDRAC stays where it is; it is the reinstall path, and192.168.90.248answers Redfish. - Use a static IP in the PVE answer file. The
.248DHCP lease is bound to a MAC that is not the active port's, and SFP+ negotiates too slowly for installer DHCP anyway (same trap as rumba). - The 8-bay chassis is why
pbs-02does not fit this machine — moot since 2026-08-20: the backup box is a bought R740xd, and this machine's job ismgmt-01.
Plan
- Gate: confirm with the K8s course owners (artifacts belong to
pim/remi) that the cluster is not needed for the autumn 2026 semester. - Rescue (total well under 100 MB of unique data):
/home/pmudry/rumba-preinstall-backup(7 MB) — the only irreplaceable data on the box (old rumba/home+/etcreferenced in the rumba history page) → copy to rumba/hdd/backupor the NAS.- Tarball of
/etc(etckeeper git repo with daily autocommits = full config history). - Skippable: pim's 46-line uncommitted playbook diff (concepts already in
ansible-playbooks-conf), registry blobs, the 2024/backup_calypso-master_conf, etcd backups, the five playbook copies (all preserved inansible-playbooks-archiveon GitHub).
- Reinstall Proxmox VE over the iDRAC with the proven remote-install process — build the ISO elsewhere (the process notes suggest calypsomaster itself, which won't exist mid-reinstall).
- Local storage: H730 into HBA mode, ZFS on the 4 × Toshiba 1.92 TB SATA SSD already in the bays. The PM1735 cards are all spoken for by the EPYC nodes — none land here.
- Standalone host, not a cluster member — it carries the guests that must survive a cluster outage (PDM, NOC, SOC, MAAS) and the corosync QDevice that gives the four-node production cluster its fifth vote.
- Cleanup: router DNS/lease rename (
calypsomaster,calypso-master,.isc3entries), move the host out of themgmtgroup inansible-playbooks-conf'sisc3_rack.yml, purge kubelet remnants oncalypso9/calypso10when those are next reinstalled. - Docs: calypso-stack.md (registry section is wrong
today regardless), incidents.md (recovery tunnel →
rumba), thermal-protection.md (orchestrator role →
srv-status), rack.mdx (role/name), tooling/ansible.md (pre-split layout, see the Various item above).
Infrastructure as code — gaps found 2026-08-04
Audit of provisioning/ against the live rack, asking whether it alone could recreate the
Proxmox services. It could not: it covered the services well but not the substrate, and two
services' real configuration existed only inside a database. Four of the original seven items are
closed (see the journal); the three below are what remains.
What is already sound, and needs no work: guests 100–104 and 108–110 are each created and
configured by their own idempotent deploy script, secrets are generated in-container rather than
committed, and srv-web01's live Caddyfile is byte-identical to the repo copy. VM 107 (a DR
artifact with a tested runbook) is deliberately unscripted.
Standing rule from the closed items: after any change in the Keycloak admin console or the NetBird dashboard, re-run that service's export script and commit the diff — otherwise the change exists nowhere but a database.
- [HIGH]
srv-pbshas no provisioning directory — and it is the guest holding the backups. Its answer file exists in one place only, on rumba, and carries a secret, so it cannot be committed as-is; the datastore, the prune and verify jobs and the token-ACL split are prose on the service page. A dead NAS currently takes the only copy of that configuration with it. - [HIGH] The FS2500's own configuration is captured nowhere — shares, NFS exports, users, the VMM
guest. The MikroTik
/exportfiles undersecretzone/mikrotik/are the model to copy (current as of 2026-08-03; restorable, just not runnable). Public DNS at Infomaniak is likewise prose only.
Email
Outbound mail works since 2026-08-10, through srv-mail (CT 112), the Postfix relay that holds
the only copy of the Infomaniak credential. A new consumer needs no password: point it at
srv-mail.isc3:25 and add its address to MYNETWORKS in
provisioning/mail/deploy-relay.sh. Design, the scope of the SInf egress allow and the
access-control layers: Email.
Wired up and verified: PVE notifications, rumba's host mail (ZFS zed, smartd, cron, UPS — zed
and smartd had been mailing into a void for months), PBS, the PDU
thermal alarm, rumba's iDRAC9 and the FS2500. Recipients are one line, ADMINS in
provisioning/mail/deploy-relay.sh — see the group address.
Open:
- Two appliances still name a person — the PDU and the FS2500 carry
pierre-andre.mudry@hevs.chin their own recipient field rather than the group address, so a change toADMINSdoes not reach them. The PDU needspdu-email.shfrom the Mac (expectexists nowhere else); the NAS recipient is UI-only, Control Panel → Notification → Email. - An alert path that does not depend on the rack — the relay,
srv-statusand everything else run on rumba, so one rumba outage or a thermal event silences every channel at once. Candidates: a channel via hannibal, or an external heartbeat alerting on absence of signal. This is the remaining gap in alerting. - Keycloak —
smtpServeris empty. Decide first whether an SSO-only realm should send any mail at all; probably not. - MikroTiks (
/tool e-mail) — deferred. Letting the CCR2004 send means putting192.168.88.1in the allowlist, which is the address all legacy WireGuard traffic is NATed to; the CRS326 at.254has no such problem. Low value either way.
Standing constraint: the allow is to Infomaniak only (465 and, since 2026-08-11, 587), so a fallback can only ever be another port to the same provider, not a second one, and Telegram stays the primary alert channel.
Utility tools
Open items for the three browser tools deployed 2026-08-21. All three are live on their public names.
- Announce them — the isc-hub cards are in place, in both the students and the teaching-staff universes (section 05).
- Decide whether the tools carry an upgrade schedule — both deploy scripts install the latest upstream release when re-run, so an upgrade is one command; nothing schedules it.
ISC Learn
Open items for the ISC Learn Moodle and its server hannibal. Finished work lands in the Learn history.
Decisions
- [HIGH] Decide and run the
migration to managed hosting. Phase-0 items to start
regardless of timing: confirm the Jelastic Cloud pérennité with Infomaniak (both the VPL jail and
the Jobe server live there), get a tier/cost quote (~8 CPU / 24 GB / 500 GB — the current VPS cost
is undocumented), confirm the direct-OIDC route with SWITCH edu-ID, ask SInf to retarget the
isc.hevs.chCNAME tolearn.isc-vs.ch(no-op today), pick the window. The managed server has to serve the hub at the root as well (plan §3). - [HIGH] Give the DRP a PDF and paper form, so the instructions survive hannibal, this site and the HES network being down at once (Remi's papers). The offline copy is what the break-glass paper key already assumes. The page was rewritten as a decision table on 2026-08-19; what it still lacks is the offline form.
- Finish the learn playbook, so hannibal can be recreated from scratch quickly (Remi's papers). The rumba copy holds the whole host config since 2026-09-05; what is missing is the playbook that replays it.
- [MEDIUM] Monitoring for ISC Learn (Remi's papers): decide where it runs — rumba is the candidate — then build the dashboard and wire its alerts through the rack relay.
Backups
- [HIGH] DS923↔rack replication, so ISC Learn backups get a
second reachable copy and the restore path can be tested. The landing spot needs a new decision:
FS2500 Volume 2 was the candidate and is now the PBS datastore,
leaving rumba's
hddpool or a share on Volume 1. The DS923 is unreachable from the rack, so the replication has to run from the school-intranet side — backups, uplink restrictions. - Schedule the rumba pull (Remi's papers): the DS923 rsync is
--delete-beforeinto one directory (versions, if any, are Synology snapshots nobody has verified); the rumba copy is versioned by ZFS snapshots since 2026-09-05 but still pulled by hand through a tunnel from an admin Mac (refresh procedure). - Finish the rsync-with-ACLs test (Remi's papers), then replace the
hannibal.sh/marcellus.shscripts on the desktop NAS and set the same up on rumba's — either straight from the source at another time of night, or as a copy of the first backup to spare the production I/O.
hannibal — follow-ups of the 2026-09-07 hardening
Log; standing state on the server page.
- Reboot for kernel 6.8.0-139 and libc6:
provisioning/hannibal/system-update.sh --reboot, about two minutes of Learn downtime. Pending since 2026-09-07. - Restrict port 20002 at the Infomaniak firewall to the HES ranges and the admins' home addresses — 1 200 failed attempts a day; the host has no firewall of its own (audit finding F-OPS-3, the Infomaniak filter now documented as the first layer).
marks_dev/marks_prod: the Marks-crawler containers run, but no key is left on the accounts. Either someone takes the service over (add their key) or the containers and the two accounts go.- Move
ingegamez.isc-vs.choff hannibal to a guest on ISC³ — it shareswww-dataand the PHP-FPM pool with Moodle, its three admins include a gmail account, andcode-snippetsgives them PHP execution. Until then: delete the inactive plugins (Elementor, Akismet) and the five unused themes, replace the gmail admin by a HES account, consider droppingcode-snippets. - Snipe-IT was deployed from leny's home and used keys of two GitHub accounts. The compose project
now lives in
/srv/docker/inventory; find out who administers the inventory today. Moving the service to the rack is under hosting. - hannibal's own sendmail — its
www-datacron output goes direct-to-MXes and bounces with550 rejected by DMARC policy, so cron errors on Moodle prod report to nobody. Outside the rack, so it cannot use the relay: give it the mailbox credentials directly.
Dated cleanups
- Delete
/srv/www/wiki.isc-vs.chon hannibal (kept a month from 2026-08-06 as the rollback path for the retired wiki; the archive is what survives). - Around 2026-09-12: delete the btrfs snapshots
/srv/.snapshots/pre-purge-2026-09-05andpre-upgrade-52-2026-09-05on hannibal (sudo btrfs subvolume delete …) — the first pins the ~92 GB the course-backup purge freed. Until then the rumba copy also still holds every deleted file; alearn-mirror-refresh.shrun propagates the deletion. - From 2026-09-12 (a week on Moodle 5.2): remove
/srv/www/learn.isc-vs.ch/moodle_isc.50(the 5.0.1 code, 452 MB) and/srv/upgrade-52/on hannibal, the same two on VM 107, and the PVE snapshotpre-upgrade-52of VM 107. Also deletemoodle_data/sessions/*on hannibal (36 MB of file sessions, unused since sessions moved to Redis). The year-move-up scripts that sat inmoodle_isc.50are kept in ISC-HEI/moodle-isc-admin-scripts (upgrade log). - From 2026-09-13 (a week of the new frontpage): delete the PVE snapshots
pre-frontpage-2aandpre-refresh-2026-09-06of VM 107. The nextlearn-mirror-refresh.shreimports the prod database, which holds the design, so the mirror converges on its own.
Moodle
- Spring 2027 registration run (February,
--semester=S2,S4,S6) per the process; before that, settle what the autumn run left out: the bachelor thesis 330.1 (22 students, labelled S6 in the workbook), threeX?marks, and whether 100.1/100.3/100.4 and 205.4 should get a Learn course. - Plan the Moodle 5.3 LTS upgrade for a holiday after its 2026-10-05 release (5.2 is supported
until 2027-04-05). Same procedure: rehearse on VM 107, then
provisioning/learn/moodle-upgrade-cutover-2026-09-05.shadapted. - Report the New Learning 12.2.7 page-width bug to the theme's author (Mariusz Boloz; the
analysis) and drop
provisioning/learn/mb2nl-fxwidth-patch.shonce a release fixes it. Until then, re-run the patch after every theme update. - Configure Moodle's router (Moodle 5.1+ feature: URLs without
r.php;admin/cli/checks.phpreportscore_routeras not configured). Not needed for the site to work; needs Apache rewrite rules per the Moodle docs. Test on VM 107 first. - New Learning's stylesheet is 1.7 MB decompressed and the browser runs no body script before it has loaded: about 4 s at 1.6 Mbit/s, now the floor under the mobile numbers (performance). Look at what the theme compiles in and whether unused parts can be left out.
- Frontpage follow-up (design): photos for a hero carousel (the hero is built for it, variant « plein cadre » in the 2026-09-06 proposals).
- Photo of Andrea Guerrieri for the teaching-team page (
mod/page/view.php?id=4067, builder page 16): upload it asfiles.isc-vs.ch/isc-learn-static/teachers/andrea.webp(portrait, same framing as the others) and put that URL in the card, which shows the petal mark meanwhile — the card is shorter than its neighbours until then.
Hosting and migrations
What still runs outside the rack, and what it takes to bring it in. ISC Learn and its server hannibal have their own section; marcellus is on the External VPS page.
- Move
isc-inventoryto a VM on ISC3, back from hannibal — the containers are identified (Aug 2026, hannibal): Snipe-IT app + MariaDB on 8080/8443. Also unblocks the learn migration's co-tenant list. The documentation already moved to Services (on-site) on 2026-08-18. - [CRITICAL]
marcellusis compromised and cut off by Infomaniak since 2026-09-07 (incident). It is discarded, not cleaned; in order:- Get the network back: send Infomaniak the post-mortem (
secretzone/marcellus-postmortem-2026-09-07.md) with the unblock request. Keep/root/quarantine-2026-09-07/until they confirm they do not want the samples. - Rebuild ikarus, do not restore it: fresh WordPress on the new host, database re-imported after its
wp_userstable has been reviewed for accounts added since 2025-11, the admin and MySQL passwords rotated — the 2021 dump, admin hashes included, was public for five years (no user accounts on that site). Ask the site owner whether the site is still wanted at all. - Treat inf1 and advpro as exposed: same
www-data, so a webshell in ikarus reached them. Rotate their admin and MySQL passwords, review their admin users and plugins before moving them anywhere. - Name the owners of the remaining vhosts (LoRa/ChirpStack, scala, sin, apt, Grafana) and of the accounts
local,tonio,tisc_*(secretzone/marcellus.md); per-vhost usage cannot be measured from the logs. Whatever nobody claims by the rebuild date is not rebuilt. TISC editor is already resolved — migrated totisc.isc-vs.chon the rack, 2026-09-18. - Move what is claimed to new guests on ISC³ or a fresh VPS (TISC editor done, 2026-09-18), then delete marcellus. The five orphan MySQL databases from the frozen Moodle removal and the
ubuntucrontab's 5-minute beacon toonline.oouu.ch(now resolving to a Brazilian address) die with the host. - Rotate the tisc-editor secrets (
AUTH_SECRET,AUTH_KEYCLOAK_SECRET,DB_PASSWORDin/root/tisc-editor/.env) — printed to a terminal while diagnosing theAUTH_URLredirect bug (2026-09-18); not otherwise exposed, but cheap to rotate.
- Get the network back: send Infomaniak the post-mortem (
Network fabric
The uplink, the switches and the node NICs — decisions that only make sense taken together.
-
[HIGH] The CCR2004 admin password in
secretzone/isc3.mdis stale — rejected by both routers (noted in the secretzone, undated). Blocks any password-based admin session, including the MikroTik backup refresh below and any emergency console access. Find or reset the current password and update the secretzone. -
Refresh the MikroTik backup (
secretzone/mikrotik/ccr2004-192.168.88.1.rsc, procedure) — stale since 2026-08-09, and now missing the srv-tisc-editor DHCP reservation and internal DNS entries added 2026-09-18. Needs the password above first. -
FS S3600-48T4S (
192.168.88.2): recover or reset the admin credentials — nothing is recorded in the secretzone, so the switch is neither backed up by Oxidized nor manageable. Once creds exist, add it toprovisioning/oxidized/router.db(oxidized has anfsosmodel). -
Add the FS switch to the stack so every iDRAC and host cable can reach every Calypso server (Remi's papers) — depends on the credentials above.
-
10 G → 25 G core? (Remi's papers) The NAS can only do 2 × 10 GbE, so the case rests on the node side, not on storage.
-
Decide
epyc1's data path. It runs on a single 1 GbE link while its pool measures 4.7–7.8 GB/s, and it holds an unused ConnectX-6 with no transceiver, in InfiniBand mode, in a rack with no 100 GbE port. Candidates: the idle second I350 port, an SFP+ NIC into the FS S3600-48T4S, or bringing the ConnectX-6 into service as part of the fabric decision. The slot contention with storage is gone: abandoning the 24-bay path frees both x16 and all four x8 FHHL slots, so four NVMe cards and the ConnectX-6 fit together. -
For the production cluster, is HDR100 enough or does HDR200 make sense, given that
epyc0andepyc1will be on it? Starting point:epyc1already carries a single-port ConnectX-6 (MT4123) on a Gen4 x16 link, which has the PCIe headroom for HDR200 (details). Check whatepyc0holds when it comes back from repair. -
Decide what the fitted ConnectX-6 is for. The card is in the machine and uncabled (found 2026-08-13); §6.1 of the target architecture plans 25 G SFP28 for
pve-03without knowing it exists. It sits in a Gen3 x8 slot, so it caps near 63 Gb/s until moved to slot 2, 5 or 7 — see PCIe slots.
Disaster recovery & documentation
- [HIGH] No rebuild runbook for rumba — the DRP covers only
ISC Learn. Much of it exists already as
post-install.shplus the per-service deploy scripts; it needs assembling and one dry read-through.
Server room
- Make something nice there — posters on the walls, screens.
- Rename the networking lab room, and change Darko's deputy at the same time.
- Install the bought patch panel outside the rack, not inside: space will be premium (purchase logged in the journal).
- Choose the new R630 and R730 / R740 for the main Rumba. Budget CHF 3 000 — reconcile with the GPU-server line in Various and with the scenario buy lists before ordering.
The server box in the networking lab room stays: one is needed for a support return, and it hides the CTF network setup used in the network labs. The fibre-speed question belongs with the fabric decisions.