Skip to main content

Open points / TODOs

Finished items move to the Ops journal so this page only holds live work.

Severity markers

[CRITICAL] / [HIGH] / [MEDIUM] prefix any item with a security or availability implication (shown as a colored badge), so a cosmetic wish is never mistaken for a data-loss risk. Unmarked items have neither. Severity is about consequence, not effort — several [CRITICAL] items are one command. The markers are still plain text under the color: search the page for [CRITICAL] to get that view.

Security audit — apply the findings

[CRITICAL] — the highest-priority item on this page. A full read-only security audit of the architecture was run on 2026-08-06. Apply its propositions. The findings, the evidence and the ordered remediation list live once, in secretzone/security-audit-2026-08.md — kept out of the published site because it names live weaknesses on named hosts.

Its "Fix today" block is down to one item: the CIFS sudoers rule (F-CODE-1). The fix is written and waits as PR branch fix/cifs-sudoers-pin (2026-08-28) — the role only reaches the CALC machines, so the merge is Gregory's call. The rest of the remediation list is still open; several of its findings already had an item on this page and are marked in place below.

Backups

  1. Wire the NAS to Healthchecks: the rack-side jobs are wired (2026-08-16, see the service page), the DSM scheduled tasks (scrubs, Hyper Backup if any) are not — needs a curl task in the DSM task scheduler.
  2. [CRITICAL] Protect the 428 GB of student homes on the FS2500isc3_homes is still the only copy: no snapshots, no off-box backup. Deferred 2026-08-02 until the Proxmox storage design settles, since rumba's hdd pool is the natural target. Two constraints when it restarts: isc3_homes is a real Btrfs subvolume, so DSM 7.2 daily and immutable snapshots are available once Snapshot Replication is installed; and it must be a push — the homes are 750 owner-only and IscAdmin cannot read them, so an rsync pulled from rumba would silently back up nothing.
  3. [HIGH] Commission the new pbs-01 — the chassis question is decided (2026-08-20, retained architecture): one R740xd all-NVMe in 307, delivered end of August 2026, 8 × PM983 RAIDZ2, no pull pair. Next: build the datastore, verify the DatastoreBackup-only token can neither prune nor delete (architecture §5), wire it as NUT master. This is the off-site leg of 3-2-1, so it lands in phase 2, before the rack move.
  4. [HIGH] Tighten the PBS credentials and token scope — see the security audit for what needs moving and rotating.
  5. [MEDIUM] Schedule a restore drill, quarterly: one CT and one VM from PBS to a scratch VMID, timed and recorded. State an RTO and RPO per service while doing it.
  6. [MEDIUM] Write the retention and deletion policy (Remi's papers): which data is kept how long, including what happens to a departing student's home. This is also the nLPD gap the audit found on the data-protection side.
  7. Create backup space for the teachers' laptops (Remi's papers).
  8. Back up rumba's own services (Remi's papers) — the guests are covered by PBS, the host's service state is not.

Various

  • Confirm the NFS export mode on the device. nas.md says async on the strength of two dated internal notes; sudo cat /etc/exports needs a password nobody has entered since. Check on the next NAS session.

  • A DNS server off the router? Today every name change means editing the CCR2004 (how it works now). Worth it only if the churn justifies a second thing to run and back up — decide against the alias-domain work already done, which removed much of the churn.

  • Order the Gigabyte G242-Z11 GPU server and its GPUs — the retained plan (2026-08-20) makes it gpu-01, the fourth production node. Specced 2026-09-01: the EPYC 7443 listing, 3 × RTX PRO 4500, 8 × 32 GB RDIMM — ≈ €14 000 (architecture §9). The card's stock is down to 3 units, so the count is also what is orderable today. The seller confirmed the chassis will come with 12VHPWR cables rather than the usual 8-pin ones. No budget is attached yet.

  • Fix the iDRAC state of the existing machines — say which and what. The known cases are the seven R630 spares whose iDRACs were never put on the management LAN (carnaval), and epyc0's BMC, which needs a reservation and DNS records when it returns. If something else is meant, name it.

  • The R7515 — pve-03, on the bench since 29 August 2026. 256 GB fitted from the epyc0 harvest, backplane cabling verified, PVE installed from a USB key (architecture §3). What is left:

    1. Six more PM983 — six are in, giving three mirror vdevs; nvme_pool reaches its 10.5 TiB with twelve, in bays 12–23.
    2. Create bulk — 10 × HPE 1.8 TB in bays 2–11, two RAIDZ2(5) vdevs. The other two carry rpool; the 300 GB SAS pair the order sheet promised was not delivered, so the boot mirror costs two of the 1.8 TB and leaves 4 spares instead of 6.
    3. The Mellanox in slot 4 comes up as InfiniBandmlxconfig … LINK_TYPE_P1=2 to put it in Ethernet mode, as on the other nodes. Model still to identify.
    4. iDRAC housekeeping — the clock reads 1998, which points at a flat CMOS battery: replace it while the lid is off, then set NTP. Address is a bench DHCP lease today and needs a rack reservation.
  • The PDU answers ICMP but not SNMP, HTTP or HTTPS from a VPN client, while an SNMP read from inside 192.168.88.0/24 works — both tested 2026-08-13, from the Mac and from epyc1. In-rack pollers such as the status page are unaffected, so this only breaks management from a laptop. Decide whether VPN clients should reach PDU management, then either open the path or record it as intended.

  • [CRITICAL] NFS configuration issue for students — say which one. This item carries a critical badge and no content. The two candidates already written down are the UID/GID mismatch that leaves five students unable to write their own ~/nas_home, and the mount re-point to 192.168.91.250. Either point this at one of them or describe the third thing it means.

  • Status-page follow-ups: create a page for upcoming scheduled events (maintenance, outages) linked from the status page, and decide whether a public services-uptime page (status.eduid.ch-style, per service rather than per host) is wanted on top of it.

  • A restricted zone for the ISC admin tooling from git, for two people only — say what "zone" means before building it: a NetBird group and policy, a Proxmox pool, a separate repository, or all three. It touches NetBird and Proxmox either way.

  • Set the OS hostnames of the Tango pair — the docs, the Ansible inventory and the status poller say tango0 / tango1; the racked node reports the factory name spark-4425.

  • Give the racked Tango node its documented 192.168.88.20 — it sits on a dynamic 192.168.88.43 because its live MAC 4C:BB:47:2F:44:25 has no reservation, while .20 / .21 are reserved for two other MACs. Decide which port is meant to be the LAN one, then either recable or move the reservation; the status poller and Ansible inventory both expect .20.

  • Review Change management and Incident management pages, the information is odd — and write the notification templates they refer to. Nothing was lost in the wiki migration: those were red links there too, so the templates have to be written from scratch; the dangling link stubs were removed 2026-08-19.

  • Network follow-ups from the 2026-08-02 audit (applied fixes in the journal, details): 0. .calypso DNS domain retired 2026-08-28 (.isc3 only, legacy records purged, node FQDNs rewritten). Left: reissue kubeconfigs that name k8s.calypso (the API certificate already lists k8s.isc3); delete the NAS snapshot /volume1/@isc3_homes-pre-rename after a few days. The guests' /etc/resolv.conf were realigned 2026-09-14 (provisioning/pve/guest-searchdomain.sh, and srv-pbs by hand); left: srv-test1 and srv-test2, stopped that day, so re-run the script once they are up. VM 107's [special:cloudinit] cache still holds searchdomain: calypso — its active config says isc3, so a cloud-init drive regeneration clears it.

    1. Fast-path route + nas-fastpath.service — the node half went with the 2026-08-28 rebuild (lab VMs mount .91.250 directly). Left: remove the unit from the Ansible role, and decide whether the NAS-side rc.d/nas-fastpath.sh route is still needed now that nothing in 192.168.91.0/24 talks to .88.250.

    2. CRS326 — the hygiene pass and the v6 → v7.23.3 upgrade are done (2026-08-04 → 09, what happened). Still open:

      1. [MEDIUM] L3 hardware offloading is now within reach: the 98DX3236 supports it under v7, and it is exactly the software-routing ceiling the 2026-08-02 audit measured (~39 MB/s against 110). Not free — inter-VLAN offload needs a vlan-filtering bridge with VLAN interfaces, where today all three subnets share one flat bridge, and offloaded traffic bypasses firewall and NAT.
      2. The target architecture still argues for retiring this switch on the grounds of "RouterOS 6.49.19, EOL, telnet enabled" — both of which are now false. The case for replacing it needs restating on current facts.
  • Pilot the exam VDI proposal (SEB + Guacamole on carnaval) with one class before any real exam — measure guacd load and per-session bandwidth against the sizing figures; the other open questions are listed on the page.

  • Open points specific to the research infrastructure (Chacha, Disco, Mambo, "Dance New") live on the CALC@HEI todo page.

EPYC nodes (epyc0 / epyc1)

Open items for the EPYC pairepyc1 is racked and running Proxmox, epyc0 is away in repair. Done work is in the journal and the history page.

  • [HIGH] Change both EPYC BMC passwords — they are the board serial numbers, printed on the chassis pull-tab and readable by any host root via dmidecode. Host root therefore escalates to BMC admin (IPMI, KVM, power) on a machine whose IPMI port every VPN user can reach. Record the replacements in secretzone/epyc.md.

  • [HIGH] The epyc1 BMC exposes IPMI (623) + Redfish on the flat 192.168.88.0/24 LAN, reachable by every VPN user — same exposure class as audit finding F-HOST-1. Fold it into whatever the flat-L2 remediation becomes.

  • [HIGH] Give epyc1 a working mail path so ZFS and PVE alerts leave the host — postfix has no relayhost and no root alias, so a degraded nvme_pool shows only in local mail and the web UI. Smarthost in secretzone/smtp.md (465 to Infomaniak only).

  • Storage plan for the EPYC nodes — decided 2026-08-13, nothing executed. Design and the disk allocation across the estate: §5.2. What is left to do:

    1. Fit the fourth PM1735 in epyc1 and rebuild nvme_pool as two mirror vdevs. The pool holds no data (1.85 MB allocated), so it costs one command — do it before anything lands on the node. Four in mirrors is larger than the current 3-wide RAIDZ1 and has no padding overhead.
    2. Four more into epyc0 when it returns from repair, same layout, same storage ID.
    3. Buy one or two 3.2 TB Gen4 NVMe spares — 4 + 4 leaves the production pair with none, and the PM1735 is out of production.
    4. Run the SFF-8654-to-U.2 cable test once (~CHF 40, SFF-8654-to-SFF-8639 with an integrated power lead). It confirms the diagnosis and keeps the repair path open; do not build a pool on it.
    5. Decide what happens to the 36 surplus PM983 — 12 spares, ~24 to sell or trade at roughly CHF 2 200–3 100 (§5.2).

    A reseller claim for the backplanes stays open as a separate option; nothing waits on it. The spare R282-Z92's untested backplane is a second one — see the item below.

  • The spare R282-Z92 — 2 × EPYC 7302, no RAM, no drives (config). Record its delivery date and invoice in the register. No CPU to buy: it boots as delivered, at 32 Zen 2 cores against a node's 48 Zen 3, and guests on cpu: x86-64-v3 run on either. What it lacks is memory — the shelf holds 2 × 32 GB, so a swap-in depends on moving the failed node's DIMMs across. Decide whether that is accepted or whether a minimal DIMM set is bought for it.

  • Test the spare R282-Z92's U.2 backplane on arrival — it is the only untested 24 × U.2 backplane in the estate, and both production ones are defective. Use the documented LED test (power only, no data cables). A good backplane makes it a fifth option alongside the reseller claim: donate it to epyc0 or epyc1 and put the PM983 in bays, or move the drives to this chassis instead. Do this before the chassis is written off as insurance only.

  • When epyc0 returns from repair: check the router for a dynamic lease rather than assuming 192.168.88.11, give its BMC a reservation and epyc0-bmc DNS records, and verify the pull-tab password in secretzone/epyc.md. Name its ZFS pool with the same storage ID as epyc1's (nvme_pool) or guests will not migrate without config edits.

  • Recover or write off the 32 GB missing on epyc1. Everything short of opening the chassis is eliminated; two physical tests remain, in order — a CMOS/NVRAM clear, then a CPU swap between sockets. Combine them with fitting the fourth PM1735 above.

  • Housekeeping on epyc1: systemctl disable zfs-import@data_pool.service retires the unit left behind by the pool deleted 2026-08-12, which fails at every boot and is the node's only failed unit; there are also 60 pending package upgrades.

  • While in epyc1's BIOS setup for anything else, consider Milan0168 Enable AER CapEnabled so PCIe link errors become visible (why). Diagnostic value only, and it would not have helped the dead bays, so it is not worth a visit of its own.

Rumba

Open items for the Proxmox node; done work is in the journal and the history page.

  • Migrate its guests to the production cluster (phase 2 of the plan). Rumba does not join the cluster and does not become mgmt-01 — that role went to calypsomaster on 2026-08-20 (iDRAC9; rumba has no BMC).
  • Define its lab/provisioning duty once emptied — decided 2026-08-20 that it stays in service on the lab side; the exact use (bench, provisioning host…) and what its 8 × 500 GB hdd pool carries are open. Its ConnectX-6 returns to the spares shelf.
  • Give isc-adm (VM 104, 192.168.88.171) an internal DNS name on the CCR2004 (isc-adm / isc-adm.isc3, per the other guests) and decide whether it gets a personal admin account instead of shared root.
  • Note before any RAM plan: all 24 DIMM slots are populated with 8 GB modules, so memory can only be replaced, never added — details.
  • Check hdd's free space before promising the pool to anything new — several items below want it as a landing spot.
  • Wire Ansible's dynamic inventory to NetBox (v2 token per consumer). The inventory itself was seeded 2026-08-16 (provisioning/netbox/populate.py); still absent from NetBox: cables/connections, the PDU outlet map, device serial numbers.

NAS (FS2500)

Ranked actions from the August 2026 disk-health readout. The 2026-08-02 hardening, first-ever data scrubbing and Volume 2 rebuild are done — history page.

  1. Re-point the node mounts to 192.168.91.250. The NAS now has one dedicated 10 GbE port per subnet — eth2 = 192.168.88.250 (DSM, rumba/PBS backups), eth3 = 192.168.91.250 static, in the node subnet (done 2026-08-03; bonding was considered and rejected — redundancy is not the goal, throughput and simplicity are).

    The node side went with the 2026-08-28 rebuild; what remains is making sure every image and lab-VM template mounts .91.250, not a per-node script. The rules and both the fstab and pvesm add nfs forms are written up once on the NAS page — how a node must mount it. No export change is needed; the NAS already exports to 192.168.91.0/24 and matches on the client address.

    Same edit drops the leftover sync and the duplicate /exports/ line from the nodes' fstab. Neither is urgent: with the exports now async, a client-side sync was measured to cost almost nothing (2026-08-05).

    Afterwards, delete /usr/local/etc/rc.d/nas-fastpath.sh, the manual 192.168.91.0/24 dev eth2 route, and nas-fastpath.service from any surviving node.

  2. Until the nodes are rebuilt, the fast-path route can vanish. Nodes still mounting 192.168.88.250 depend on it, and /usr/local/etc/rc.d/nas-fastpath.sh re-adds it only at DSM boot, so a link flap costs ~40 % of read throughput silently. Re-check ip route after physical work behind the NAS; if the rebuild is far off, a nas-fastpath.timer (OnUnitActiveSec=5min) would bound it.

  3. Tidy eth0/eth1 on the NAS. Both are still BOOTPROTO=dhcp with no cable, so plugging either one in re-opens the DSM gateway-election hole described on the NAS page.

  4. Enable DSM notifications and SSD-lifespan alerts so a degraded array or wearing disk is not discovered by accident.

  5. When the 16 GB DIMM arrives (ordered from senetic, still not delivered as of 2026-08-18): install in slot B (→ 24 GB total), then raise the PBS VM to 4 vCPU / 8 GB in VMM and eject its install ISO (Edit → Others — the API can't do it).

UPS and power

The rack UPS has powered rumba and nothing else since 2026-08-03. It is monitored with NUT and shown on the status page, which alerts on mains loss.

danger
[HIGH] NUT has been blind since 2026-08-10

The UPS dropped off rumba's USB bus and there is no remote lever: both ends were reseated, a full reboot reset the xHCI controller, neither produced a kernel event (incident). Until it is fixed a mains failure means rumba runs to a flat battery with no clean shutdown, and the status page shows stale UPS data.

Next time at the rack, in order, watching ssh root@192.168.88.51 'dmesg -TW | grep --line-buffered -iE "usb 1-|0665"':

  1. A known-good USB device in port 13 (rear panel, left, upper — where the UPS cable is). It enumerates → the cable or the UPS's own port is at fault; it does not → port 13 is dead, move the UPS to a neighbouring port.
  2. Swap the USB cable.
  3. Whatever brings it back, finish with systemctl restart nut-driver@rackups and upsc rackups; the driver is not known to reattach on its own.

1. [HIGH] Decide between repairing the USB link and moving to serial

The unit has a DB9 serial port and rumba has two real UARTs, and nutdrv_qx speaks Q* natively over serial — so the USB cable is not the only possible path, and that link has now failed three times. The cost is specific: serial carries Q* only, so it loses the HID interface and with it the genuine RunTimeToEmpty and RemainingCapacity — exactly the numbers the shutdown decision needs. Decide the two together.

2. [HIGH] Shut rumba down cleanly before the battery runs out

On a building outage the UPS carries rumba about 55 minutes, then cuts power instantly — a running hypervisor with live guests and ZFS pools. Nothing prevents that today: upsmon runs powervalue 0, deliberately. What it lacks is a battery level it can trust, since nutdrv_qx reports a percentage pinned at 100 % and the correct data sits on a USB channel NUT cannot read. Three ways to close it:

  • A. Trust NUT's LB flag — one line (powervalue 1, MINSUPPLIES 1). Unverified: nobody knows whether the firmware sets that flag or NUT synthesises it from the broken percentage. If synthesised, the shutdown fires far too early or never. Item 3 settles it.
  • B. A small shutdown watchdog of our own — a timer reads ups-read-hid and shuts down under a margin (say 15 min) or on the UPS's genuine shutdown imminent bit. ~50 lines plus a unit, and ours to maintain. Recommended: the only option acting on numbers we have verified.
  • C. Bridge the good data into NUT via a dummy-ups device kept updated by a script. Stock NUT does the deciding, but it is the most moving parts and needs a staleness guard — a dead updater leaves upsmon reading "battery full" forever.

3. [HIGH] Run one controlled power-fail test

Decided 2026-08-03, not yet run. Runbook: UPS controlled discharge test. It answers three questions at once — whether option A's LB flag is real, what the battery's autonomy actually is today (55 min is the UPS's own estimate and cells age), and whether the shutdown path works end to end. It costs a planned outage; the alternative is learning these answers during an unplanned one.

4. [MEDIUM] Decide what else the UPS should protect

rumba uses ~14 % of the 2 kVA, so there is headroom, but 2 kVA cannot cover the rack (PDU peak 3.63 kW). The obvious candidate is the NAS, which holds the only copy of the student homes. Needs a physical cord trace at the rack — and rename the PDU's 24 factory-default outlet names while there, so the question is answerable remotely next time. Also check whether the chassis has an intelligent-card slot for a real SNMP interface.

R630 hardware — fit the PERC battery kits

The five kits are on hand (pack with holder, 070K80 / 0H132V / 037CT1 / 0HD8WG). Confirmed missing (names after the 2026-08-28 renumbering): carnaval1, carnaval2, carnaval7, carnaval8; carnaval5 unconfirmed (it is also the node with the dead DIMM_B1 — same visit to the room). The per-node readings, the two hardware traps and how the fleet was swept are on the archived Calypso page.

VPN and identity — NetBird, Keycloak, edu-ID

The remote-access and single-sign-on stack: the NetBird VPN, the Keycloak broker federating to SWITCH edu-ID, and the legacy MikroTik WireGuard being retired behind them. Architecture: §7.

  • Legacy WireGuard client confs carry a dead DNS line — the 72 already-issued .conf files still say DNS = 172.30.7.1, which never resolved the internal names (the router-side client-dns was fixed 2026-08-16, history). Fix on next contact — reissue with wg-gen.sh or hand-edit to DNS = 192.168.88.1, isc3, calypso; no mass campaign, WireGuard is being retired anyway.
  • NetBird follow-ups — the deployment, the edu-ID federation (Phase 2) and the admission policy are done and validated, 2026-08-03 → 06 (journal); the current state lives on the service page and the Keycloak page. What remains:
    1. Onboard users: teachers/admins now; students at the autumn intake, who go straight to NetBird as part of retiring the legacy MikroTik WireGuard.
    2. eduPersonOrgUnitDN is not being released yetdone. The claim arrives: as of 2026-09-06 isc holds 15 members and hevs 16, students among them (*@students.hevs.ch accounts carry RACA-TICO-ISCO, which the 2026-08-05 staff-only measurement could not show). No consumer is left on hes-so: the user gate and the realm's admission gate both test isc.
    3. Verify replay-config.py --jwt-twins against the post-collapse group model — it pairs <name>/<name>-manual from the pre-2026-08-06 identities; the policies now match entitlement jwt groups (vpn-rack-operators, vpn-rack-mgmt) that have no manual twin, so a store rebuild may need it extended before the twins step works.
    4. Optional NetBird group slimming: with a single local account, one break-glass group suffices — edit teachers-manual / students-manual out of the four policies first, then delete them (deleting first cascades — incident, which is also why they exist again), keeping admins-manual; rack-router tags the routing peer but nothing references it — confirm and drop. hes-so / role-rack-admins stay: the claim carries them, JWT sync would just re-create them.
    5. [HIGH] Require MFA for the Keycloak admin account. /admin* no longer depends on the client's address — it is behind the admin gate and the group role-rack-admins (three members since August) — and the gate authenticates against Keycloak itself. Since 2026-08-05 the same group is also the sole Administrator grant on both Proxmox clusters, so it now carries two systems; root@pam remains the break-glass for that half. MFA on it would be requested per-flow via acr-values (PVE's realm takes one), which is why REFEDS MFA was deliberately left off at the edu-ID resource level.
    6. Fill roster.csv for the autumn cohort — the mechanism is done and proven with one admin, but the file holds a single line. Needs the class list and the address edu-ID actually releases for students; a wrong domain means "logs in fine, no access" for everyone. Verify with the first student login that they also carry RACA-TICO-ISCO (item 2 above).
    7. Re-verify the whole chain with a second identity. Everything so far was proven with one edu-ID account, which cannot exercise: two people with different roles at once, a user who is in no role group being refused by jwt_allow_groups (only reasoned, never observed), and the idp-auto-link path for a roster account that has never logged in. A colleague's account or the first student is the test.
    8. Rewrite roster-sync.sh on the Admin REST API (curl + jq) instead of kcadm.sh. Every kcadm call starts a JVM (2–3 s), two per rostered person, so a full pass takes minutes; the --only <email> mode added 2026-09-03 brings one arrival down to about 30 s but the yearly move-up and removals still need the full pass. One token request then plain HTTP calls would put the whole file under 15 s. Same behaviour to keep: idempotent, never deletes an account, role-rack-admins add-only.
    9. Route tango0 / tango1 to vpn-carnaval — they sit in the 88 subnet, so a host resource per machine in net-carnaval-guests is the way, and the racked node still runs on a dynamic .43 lease (tango): give it its static address first. Decided 2026-09-08: no rule until then, moving the machines physically is an option.

UID allocation — wire the register up

The register exists since 2026-08-05: provisioning/uid/uid-map.csv, seeded from home ownership on the NAS, issued by uid-alloc.py and made into a home by create-home.sh. Ansible reads it on the rack; what is left is applying it and deciding about the machines that share the same lists.

  1. The 13 renumberings are moot — the /etc/passwd carrying the drift was wiped with the 2026-08-28 node rebuild. Make the lab-VM templates draw their accounts from the register — that is the item that matters.
  2. The research machines are frozen — Gregory has two PRs open on dance.yml. Do not touch that inventory or uid_others until they land. Afterwards: the conf/users id: columns survive only because those lists serve both fleets with different numbering, so deleting them means unifying the two namespaces first — renumber and chown one side.
  3. Point carnaval-lab-vm.sh at the register rather than stating the NAS. Not urgent: the NAS ownership is what NFS enforces, so it cannot be stale — but it means two code paths for one fact. This one is the code path that stays, so it is the item that matters.
  4. Re-run uid-alloc.py --verify whenever the NAS homes change. The nodes that were down when the register was seeded (calypso37) were wiped before coming back, so their local accounts never need sweeping.
  5. Simplify 01_users.yml's register logic once louis.heredero is unified (see the Carnaval item 7 fix). The canon/solo/status bookkeeping in provisioning/ansible/roles/00_setup/tasks/01_users.yml — one CSV row per number, canonical only when a name has exactly one non-reserved row — exists solely to represent a name split across two live numbers on two fleets. Once he's down to one number, that becomes the only case it was handling; collapse it back to a plain name → uid dict.

Worth doing before the autumn intake: that is the next batch of numbers to hand out, and the point at which pre-created roster accounts start needing them.

Carnaval (playground cluster)

Ten-node cluster since 2026-08-28 — the cluster page holds the current state. What remains:

  1. carnaval5 — replace or pull DIMM_B1, then rebuild it (perc-raid1.sh, baked ISO, post-install, pvecm addthe recipe). The missing PERC batteries (carnaval7, 8, …) are the fleet-wide sweep in R630 hardware. Also: remove the legacy calypsoN DNS records on the CCR once nothing references them, and register the ten nodes in NetBox.

  2. Alert when a fast pool suspends, and decide what happens to the Swissbit cards. Both cards of the pair have now hung the same way — carnaval1 on 2026-08-06, unnoticed for 17 h (incident), and carnaval8 on 2026-09-03, unnoticed for five days (incident). The iDRAC cannot see an add-in NVMe, so no SEL entry and no health change. Two separate jobs:

    • Configure zfs-zed to mail on zpool suspended / device removal on every node. Cheap, independent of the planned NOC, and it covers the only failure this cluster has had — twice.
    • Decide whether the cards stay in service. carnaval8 is powered off holding the one that hung in September (EUI …03550000, 7398 h); gpu9 currently runs on the one that hung in August (…05590000, 6555 h, in carnaval9); carnaval7 carries a third of the same model. SMART is clean on all of them and was clean before both hangs, so it is no help deciding — the August decision to keep the card rested on exactly that clean report. Check whether Swissbit has firmware newer than ARR50002; there is no warranty case on wear, only on the hang. A card pulled has to be replaced: these are the only 894 GB guest pools in the fleet.
  3. Kubernetes — phase 2. The shared cluster built 2026-08-28 (bash, not the playbook the design foresaw) has no VMs since 2026-09-18 (gpu02 hold the cards; carnaval-k8s.sh up rebuilds it from the current template in ~10 min). Left: edu-ID student access (k3s OIDC → Keycloak, namespace-per-cohort RBAC + ResourceQuota, blocked on item 5), per-team clusters as further instances of the same script, and per-student storage quotas (NAS and local). Student access is wanted for the last block of the Kubernetes course, around late October 2026 (lecturer, 2026-09-09); the fallback is k3s on their laptops.

  4. Move the guest-facing half to Ansible. carnaval-lab-vm.sh reads who gets in from ISC³-only files since 2026-08-16 (provisioning/keycloak/roster.csv + ssh-keys, decoupled from the Ansible tree calc shares) — and since 2026-08-05 it reads UIDs from the NAS home ownership rather than a login node, so it no longer depends on one at all. community.general.proxmox_kvm does what the script does with proper idempotence instead of hand-built cloud-init YAML. The substrate scripts (install, pool, cluster) can stay as they are. The UID columns in conf/users/*.yml should be deleted, not repaired, as part of the same job — see UID allocation.

  5. Grant the cohorts PVEVMUser on their pools once PVE accounts exist — there are none yet, so there are no student accounts yet. The federation half is done since 2026-08-05 — the cluster takes an edu-ID login through Keycloak, but only for pre-created administrators. What is left is the part that avoids hand-creating ~40 accounts a year: --groups-claim groups plus --autocreate 1, then the pool ACLs on the claim-derived students group rather than on individuals.

  6. Then decide which guests deserve a selective backup job — there is deliberately none today. The templates are covered since 2026-08-16 (publish leaves a vzdump on nas-library); the question remains for infra guests, if the cluster ever grows any worth keeping.

  7. Five students' UIDs vs GIDs — the Calypso-side half of this is gone with the nodes (2026-08-28); what remains is making sure the lab VMs issue the register's numbers (provisioning/uid/uid-map.csv) so ~/nas_home is writable — verify on the first cohort VMs.

  8. Publish the 2026-09-18 CUDA template to carnaval3, 4, 6, 7, 10 once they are powered on — they still carry the August image without Slurm/MPI. Per node: qm destroy 910N --purge 1, then NODES="carnaval9 carnavalN" provisioning/pve/carnaval-guests.sh --cuda-seal (pipeline).

  9. Collect the 25/26 cohort's SSH keys into provisioning/keycloak/ssh-keys (format). Every roster person has an account on every lab VM since 2026-08-16, but only 8 of the 34 vpn-carnaval have a key — the other 24 have an account on gpu9 with no way into it (2026-09-08). Existing VMs do not pick up roster changes, so the keys have to be in before the VM the cohort will use is built. adrien.reynard now has a NAS home; loic.christen1 needed the nasname column instead, his home being loic.christen (fixed 2026-09-08, his account appears at the next provisioning run).

Calypsomaster → mgmt-01

Audited read-only 2026-08-02 and inventoried over SSH and Redfish 2026-08-13: nothing load-bearing runs on it. Decided 2026-08-20 (retained architecture): it becomes mgmt-01, the standalone watcher outside the production cluster — PDM, NOC, SOC, MAAS and the corosync QDevice — taking the role from rumba for its iDRAC9, faster CPUs and free DIMM slots. The conversion can run now.

Nothing depends on it. Its kubeadm control plane has no user workloads and its only enrolled workers (calypso9/calypso10) are now carnaval nodes; MAAS 3.6.2 has exactly one machine in its inventory — itself, DHCP off everywhere; the Docker registry container was started without published ports, so it has been unreachable all along (5.2 GB of old course images); Zabbix Agent 2 points at a server on epyc0 that no longer exists. Nothing external depends on it either — the carnaval nodes use router DNS and mount the NAS directly. The only soft dependency is the NAS-lockout recovery tunnel in the incident log, and rumba serves that role identically.

What the conversion needs to know. The machine and its disks are specced on the archived Calypso page; the rest is:

  • The PERC H730 is hardware RAID — put it in HBA mode before ZFS touches it. U.2 drives do not fit (SAS/SATA backplane), so NVMe has to arrive on add-in cards; the PERC occupies slot 2.
  • No recabling: its 10 G DAC lands on the CCR2004 sfp-sfpplus4, same bridge-LAN as rumba — already in the admin segment. The idle 1 GbE ports (eno3/eno4) are a free option for a dedicated corosync link later. iDRAC stays where it is; it is the reinstall path, and 192.168.90.248 answers Redfish.
  • Use a static IP in the PVE answer file. The .248 DHCP lease is bound to a MAC that is not the active port's, and SFP+ negotiates too slowly for installer DHCP anyway (same trap as rumba).
  • The 8-bay chassis is why pbs-02 does not fit this machine — moot since 2026-08-20: the backup box is a bought R740xd, and this machine's job is mgmt-01.

Plan

  1. Gate: confirm with the K8s course owners (artifacts belong to pim/remi) that the cluster is not needed for the autumn 2026 semester.
  2. Rescue (total well under 100 MB of unique data):
    • /home/pmudry/rumba-preinstall-backup (7 MB) — the only irreplaceable data on the box (old rumba /home + /etc referenced in the rumba history page) → copy to rumba /hdd/backup or the NAS.
    • Tarball of /etc (etckeeper git repo with daily autocommits = full config history).
    • Skippable: pim's 46-line uncommitted playbook diff (concepts already in ansible-playbooks-conf), registry blobs, the 2024 /backup_calypso-master_conf, etcd backups, the five playbook copies (all preserved in ansible-playbooks-archive on GitHub).
  3. Reinstall Proxmox VE over the iDRAC with the proven remote-install process — build the ISO elsewhere (the process notes suggest calypsomaster itself, which won't exist mid-reinstall).
  4. Local storage: H730 into HBA mode, ZFS on the 4 × Toshiba 1.92 TB SATA SSD already in the bays. The PM1735 cards are all spoken for by the EPYC nodes — none land here.
  5. Standalone host, not a cluster member — it carries the guests that must survive a cluster outage (PDM, NOC, SOC, MAAS) and the corosync QDevice that gives the four-node production cluster its fifth vote.
  6. Cleanup: router DNS/lease rename (calypsomaster, calypso-master, .isc3 entries), move the host out of the mgmt group in ansible-playbooks-conf's isc3_rack.yml, purge kubelet remnants on calypso9/calypso10 when those are next reinstalled.
  7. Docs: calypso-stack.md (registry section is wrong today regardless), incidents.md (recovery tunnel → rumba), thermal-protection.md (orchestrator role → srv-status), rack.mdx (role/name), tooling/ansible.md (pre-split layout, see the Various item above).

Infrastructure as code — gaps found 2026-08-04

Audit of provisioning/ against the live rack, asking whether it alone could recreate the Proxmox services. It could not: it covered the services well but not the substrate, and two services' real configuration existed only inside a database. Four of the original seven items are closed (see the journal); the three below are what remains.

What is already sound, and needs no work: guests 100–104 and 108–110 are each created and configured by their own idempotent deploy script, secrets are generated in-container rather than committed, and srv-web01's live Caddyfile is byte-identical to the repo copy. VM 107 (a DR artifact with a tested runbook) is deliberately unscripted.

Standing rule from the closed items: after any change in the Keycloak admin console or the NetBird dashboard, re-run that service's export script and commit the diff — otherwise the change exists nowhere but a database.

  1. [HIGH] srv-pbs has no provisioning directory — and it is the guest holding the backups. Its answer file exists in one place only, on rumba, and carries a secret, so it cannot be committed as-is; the datastore, the prune and verify jobs and the token-ACL split are prose on the service page. A dead NAS currently takes the only copy of that configuration with it.
  2. [HIGH] The FS2500's own configuration is captured nowhere — shares, NFS exports, users, the VMM guest. The MikroTik /export files under secretzone/mikrotik/ are the model to copy (current as of 2026-08-03; restorable, just not runnable). Public DNS at Infomaniak is likewise prose only.

Email

Outbound mail works since 2026-08-10, through srv-mail (CT 112), the Postfix relay that holds the only copy of the Infomaniak credential. A new consumer needs no password: point it at srv-mail.isc3:25 and add its address to MYNETWORKS in provisioning/mail/deploy-relay.sh. Design, the scope of the SInf egress allow and the access-control layers: Email.

Wired up and verified: PVE notifications, rumba's host mail (ZFS zed, smartd, cron, UPS — zed and smartd had been mailing into a void for months), PBS, the PDU thermal alarm, rumba's iDRAC9 and the FS2500. Recipients are one line, ADMINS in provisioning/mail/deploy-relay.sh — see the group address.

Open:

  1. Two appliances still name a person — the PDU and the FS2500 carry pierre-andre.mudry@hevs.ch in their own recipient field rather than the group address, so a change to ADMINS does not reach them. The PDU needs pdu-email.sh from the Mac (expect exists nowhere else); the NAS recipient is UI-only, Control Panel → Notification → Email.
  2. An alert path that does not depend on the rack — the relay, srv-status and everything else run on rumba, so one rumba outage or a thermal event silences every channel at once. Candidates: a channel via hannibal, or an external heartbeat alerting on absence of signal. This is the remaining gap in alerting.
  3. KeycloaksmtpServer is empty. Decide first whether an SSO-only realm should send any mail at all; probably not.
  4. MikroTiks (/tool e-mail) — deferred. Letting the CCR2004 send means putting 192.168.88.1 in the allowlist, which is the address all legacy WireGuard traffic is NATed to; the CRS326 at .254 has no such problem. Low value either way.

Standing constraint: the allow is to Infomaniak only (465 and, since 2026-08-11, 587), so a fallback can only ever be another port to the same provider, not a second one, and Telegram stays the primary alert channel.

Utility tools

Open items for the three browser tools deployed 2026-08-21. All three are live on their public names.

  • Announce them — the isc-hub cards are in place, in both the students and the teaching-staff universes (section 05).
  • Decide whether the tools carry an upgrade schedule — both deploy scripts install the latest upstream release when re-run, so an upgrade is one command; nothing schedules it.

ISC Learn

Open items for the ISC Learn Moodle and its server hannibal. Finished work lands in the Learn history.

Decisions

  • [HIGH] Decide and run the migration to managed hosting. Phase-0 items to start regardless of timing: confirm the Jelastic Cloud pérennité with Infomaniak (both the VPL jail and the Jobe server live there), get a tier/cost quote (~8 CPU / 24 GB / 500 GB — the current VPS cost is undocumented), confirm the direct-OIDC route with SWITCH edu-ID, ask SInf to retarget the isc.hevs.ch CNAME to learn.isc-vs.ch (no-op today), pick the window. The managed server has to serve the hub at the root as well (plan §3).
  • [HIGH] Give the DRP a PDF and paper form, so the instructions survive hannibal, this site and the HES network being down at once (Remi's papers). The offline copy is what the break-glass paper key already assumes. The page was rewritten as a decision table on 2026-08-19; what it still lacks is the offline form.
  • Finish the learn playbook, so hannibal can be recreated from scratch quickly (Remi's papers). The rumba copy holds the whole host config since 2026-09-05; what is missing is the playbook that replays it.
  • [MEDIUM] Monitoring for ISC Learn (Remi's papers): decide where it runs — rumba is the candidate — then build the dashboard and wire its alerts through the rack relay.

Backups

  • [HIGH] DS923↔rack replication, so ISC Learn backups get a second reachable copy and the restore path can be tested. The landing spot needs a new decision: FS2500 Volume 2 was the candidate and is now the PBS datastore, leaving rumba's hdd pool or a share on Volume 1. The DS923 is unreachable from the rack, so the replication has to run from the school-intranet side — backups, uplink restrictions.
  • Schedule the rumba pull (Remi's papers): the DS923 rsync is --delete-before into one directory (versions, if any, are Synology snapshots nobody has verified); the rumba copy is versioned by ZFS snapshots since 2026-09-05 but still pulled by hand through a tunnel from an admin Mac (refresh procedure).
  • Finish the rsync-with-ACLs test (Remi's papers), then replace the hannibal.sh / marcellus.sh scripts on the desktop NAS and set the same up on rumba's — either straight from the source at another time of night, or as a copy of the first backup to spare the production I/O.

hannibal — follow-ups of the 2026-09-07 hardening

Log; standing state on the server page.

  • Reboot for kernel 6.8.0-139 and libc6: provisioning/hannibal/system-update.sh --reboot, about two minutes of Learn downtime. Pending since 2026-09-07.
  • Restrict port 20002 at the Infomaniak firewall to the HES ranges and the admins' home addresses — 1 200 failed attempts a day; the host has no firewall of its own (audit finding F-OPS-3, the Infomaniak filter now documented as the first layer).
  • marks_dev / marks_prod: the Marks-crawler containers run, but no key is left on the accounts. Either someone takes the service over (add their key) or the containers and the two accounts go.
  • Move ingegamez.isc-vs.ch off hannibal to a guest on ISC³ — it shares www-data and the PHP-FPM pool with Moodle, its three admins include a gmail account, and code-snippets gives them PHP execution. Until then: delete the inactive plugins (Elementor, Akismet) and the five unused themes, replace the gmail admin by a HES account, consider dropping code-snippets.
  • Snipe-IT was deployed from leny's home and used keys of two GitHub accounts. The compose project now lives in /srv/docker/inventory; find out who administers the inventory today. Moving the service to the rack is under hosting.
  • hannibal's own sendmail — its www-data cron output goes direct-to-MXes and bounces with 550 rejected by DMARC policy, so cron errors on Moodle prod report to nobody. Outside the rack, so it cannot use the relay: give it the mailbox credentials directly.

Dated cleanups

  • Delete /srv/www/wiki.isc-vs.ch on hannibal (kept a month from 2026-08-06 as the rollback path for the retired wiki; the archive is what survives).
  • Around 2026-09-12: delete the btrfs snapshots /srv/.snapshots/pre-purge-2026-09-05 and pre-upgrade-52-2026-09-05 on hannibal (sudo btrfs subvolume delete …) — the first pins the ~92 GB the course-backup purge freed. Until then the rumba copy also still holds every deleted file; a learn-mirror-refresh.sh run propagates the deletion.
  • From 2026-09-12 (a week on Moodle 5.2): remove /srv/www/learn.isc-vs.ch/moodle_isc.50 (the 5.0.1 code, 452 MB) and /srv/upgrade-52/ on hannibal, the same two on VM 107, and the PVE snapshot pre-upgrade-52 of VM 107. Also delete moodle_data/sessions/* on hannibal (36 MB of file sessions, unused since sessions moved to Redis). The year-move-up scripts that sat in moodle_isc.50 are kept in ISC-HEI/moodle-isc-admin-scripts (upgrade log).
  • From 2026-09-13 (a week of the new frontpage): delete the PVE snapshots pre-frontpage-2a and pre-refresh-2026-09-06 of VM 107. The next learn-mirror-refresh.sh reimports the prod database, which holds the design, so the mirror converges on its own.

Moodle

  • Spring 2027 registration run (February, --semester=S2,S4,S6) per the process; before that, settle what the autumn run left out: the bachelor thesis 330.1 (22 students, labelled S6 in the workbook), three X? marks, and whether 100.1/100.3/100.4 and 205.4 should get a Learn course.
  • Plan the Moodle 5.3 LTS upgrade for a holiday after its 2026-10-05 release (5.2 is supported until 2027-04-05). Same procedure: rehearse on VM 107, then provisioning/learn/moodle-upgrade-cutover-2026-09-05.sh adapted.
  • Report the New Learning 12.2.7 page-width bug to the theme's author (Mariusz Boloz; the analysis) and drop provisioning/learn/mb2nl-fxwidth-patch.sh once a release fixes it. Until then, re-run the patch after every theme update.
  • Configure Moodle's router (Moodle 5.1+ feature: URLs without r.php; admin/cli/checks.php reports core_router as not configured). Not needed for the site to work; needs Apache rewrite rules per the Moodle docs. Test on VM 107 first.
  • New Learning's stylesheet is 1.7 MB decompressed and the browser runs no body script before it has loaded: about 4 s at 1.6 Mbit/s, now the floor under the mobile numbers (performance). Look at what the theme compiles in and whether unused parts can be left out.
  • Frontpage follow-up (design): photos for a hero carousel (the hero is built for it, variant « plein cadre » in the 2026-09-06 proposals).
  • Photo of Andrea Guerrieri for the teaching-team page (mod/page/view.php?id=4067, builder page 16): upload it as files.isc-vs.ch/isc-learn-static/teachers/andrea.webp (portrait, same framing as the others) and put that URL in the card, which shows the petal mark meanwhile — the card is shorter than its neighbours until then.

Hosting and migrations

What still runs outside the rack, and what it takes to bring it in. ISC Learn and its server hannibal have their own section; marcellus is on the External VPS page.

  • Move isc-inventory to a VM on ISC3, back from hannibal — the containers are identified (Aug 2026, hannibal): Snipe-IT app + MariaDB on 8080/8443. Also unblocks the learn migration's co-tenant list. The documentation already moved to Services (on-site) on 2026-08-18.
  • [CRITICAL] marcellus is compromised and cut off by Infomaniak since 2026-09-07 (incident). It is discarded, not cleaned; in order:
    1. Get the network back: send Infomaniak the post-mortem (secretzone/marcellus-postmortem-2026-09-07.md) with the unblock request. Keep /root/quarantine-2026-09-07/ until they confirm they do not want the samples.
    2. Rebuild ikarus, do not restore it: fresh WordPress on the new host, database re-imported after its wp_users table has been reviewed for accounts added since 2025-11, the admin and MySQL passwords rotated — the 2021 dump, admin hashes included, was public for five years (no user accounts on that site). Ask the site owner whether the site is still wanted at all.
    3. Treat inf1 and advpro as exposed: same www-data, so a webshell in ikarus reached them. Rotate their admin and MySQL passwords, review their admin users and plugins before moving them anywhere.
    4. Name the owners of the remaining vhosts (LoRa/ChirpStack, scala, sin, apt, Grafana) and of the accounts local, tonio, tisc_* (secretzone/marcellus.md); per-vhost usage cannot be measured from the logs. Whatever nobody claims by the rebuild date is not rebuilt. TISC editor is already resolved — migrated to tisc.isc-vs.ch on the rack, 2026-09-18.
    5. Move what is claimed to new guests on ISC³ or a fresh VPS (TISC editor done, 2026-09-18), then delete marcellus. The five orphan MySQL databases from the frozen Moodle removal and the ubuntu crontab's 5-minute beacon to online.oouu.ch (now resolving to a Brazilian address) die with the host.
    6. Rotate the tisc-editor secrets (AUTH_SECRET, AUTH_KEYCLOAK_SECRET, DB_PASSWORD in /root/tisc-editor/.env) — printed to a terminal while diagnosing the AUTH_URL redirect bug (2026-09-18); not otherwise exposed, but cheap to rotate.

Network fabric

The uplink, the switches and the node NICs — decisions that only make sense taken together.

  • [HIGH] The CCR2004 admin password in secretzone/isc3.md is stale — rejected by both routers (noted in the secretzone, undated). Blocks any password-based admin session, including the MikroTik backup refresh below and any emergency console access. Find or reset the current password and update the secretzone.

  • Refresh the MikroTik backup (secretzone/mikrotik/ccr2004-192.168.88.1.rsc, procedure) — stale since 2026-08-09, and now missing the srv-tisc-editor DHCP reservation and internal DNS entries added 2026-09-18. Needs the password above first.

  • FS S3600-48T4S (192.168.88.2): recover or reset the admin credentials — nothing is recorded in the secretzone, so the switch is neither backed up by Oxidized nor manageable. Once creds exist, add it to provisioning/oxidized/router.db (oxidized has an fsos model).

  • Add the FS switch to the stack so every iDRAC and host cable can reach every Calypso server (Remi's papers) — depends on the credentials above.

  • 10 G → 25 G core? (Remi's papers) The NAS can only do 2 × 10 GbE, so the case rests on the node side, not on storage.

  • Decide epyc1's data path. It runs on a single 1 GbE link while its pool measures 4.7–7.8 GB/s, and it holds an unused ConnectX-6 with no transceiver, in InfiniBand mode, in a rack with no 100 GbE port. Candidates: the idle second I350 port, an SFP+ NIC into the FS S3600-48T4S, or bringing the ConnectX-6 into service as part of the fabric decision. The slot contention with storage is gone: abandoning the 24-bay path frees both x16 and all four x8 FHHL slots, so four NVMe cards and the ConnectX-6 fit together.

  • For the production cluster, is HDR100 enough or does HDR200 make sense, given that epyc0 and epyc1 will be on it? Starting point: epyc1 already carries a single-port ConnectX-6 (MT4123) on a Gen4 x16 link, which has the PCIe headroom for HDR200 (details). Check what epyc0 holds when it comes back from repair.

  • Decide what the fitted ConnectX-6 is for. The card is in the machine and uncabled (found 2026-08-13); §6.1 of the target architecture plans 25 G SFP28 for pve-03 without knowing it exists. It sits in a Gen3 x8 slot, so it caps near 63 Gb/s until moved to slot 2, 5 or 7 — see PCIe slots.

Disaster recovery & documentation

  • [HIGH] No rebuild runbook for rumba — the DRP covers only ISC Learn. Much of it exists already as post-install.sh plus the per-service deploy scripts; it needs assembling and one dry read-through.

Server room

  1. Make something nice there — posters on the walls, screens.
  2. Rename the networking lab room, and change Darko's deputy at the same time.
  3. Install the bought patch panel outside the rack, not inside: space will be premium (purchase logged in the journal).
  4. Choose the new R630 and R730 / R740 for the main Rumba. Budget CHF 3 000 — reconcile with the GPU-server line in Various and with the scenario buy lists before ordering.

The server box in the networking lab room stays: one is needed for a support return, and it hides the CTF network setup used in the network labs. The fibre-speed question belongs with the fabric decisions.