Skip to main content

Ops journal

Completed actions, newest first. This is where finished items from the todo land, so that page only ever contains live work. Each entry links to the page holding the details.

2026-09-21

  • phpIPAM deployment - srv-phpipam deployed on VM 124 rumba, published at http://srv-phpipam.isc3. TLS has not been implemented yet. Service discovery is performed by a cron job running every 30 minutes. Documentation is available here.
  • phpIPAM - mailing - Add phpIPAM ip in the networks list in mail system

2026-09-18

  • TISC editor migrated off marcellus to its own guestsrv-tisc-editor (LXC 123, rumba), published at https://tisc.isc-vs.ch through srv-web01; DHCP reservation + internal DNS added on the CCR2004, public CNAME at Infomaniak, marcellus.isc-vs.ch DNS removed (service, history). A baked-in AUTH_URL pointing at the guest's private IP made every login redirect unreachable from outside the VPN; fixed the same day. Marcellus decommission list: ops todo → hosting.
  • Slurm and MPI on the GPU VMs — one scheduler per VM, the card as gres/gpu:1, OpenMPI over PMIx, in the CUDA recipe (Slurm). Template rebuilt and published to the five online nodes, k8s02 destroyed on the way, gpu02 and gpu9 re-stamped (history); the five offline nodes are in the todo.
  • gpu0gpu2 replace the k8s VMs on carnaval02 — k3s servers stopped (onboot 0), then destroyed the same day for the template rebuild; three GPU lab VMs stamped with the roster (history). Kubernetes phase 2 stays open (ops todo → Carnaval).
  • The ubuntu account's keys come from the rosterrole-rack-admins × ssh-keys via provisioning/pve/admin-keys.sh in the three carnaval scripts, pushed to the running guests (account model).

2026-09-14

  • Rack mail fans out to a group address — every consumer now writes to rack-admins@isc-vs.ch, expanded on srv-mail, so who gets alerted is one line (ADMINS in provisioning/mail/deploy-relay.sh) instead of a field per device; Yacine added (group address). Applied and verified on the relay, rumba's root alias and PVE endpoint, PBS and the iDRAC9 — the PDU and the NAS still name a person (ops todo → Email).
  • PBS notifications had never been able to send — the rack-relay endpoint carried srv-mail.calypso, NXDOMAIN from that VM. Corrected to srv-mail.isc3 by provisioning/mail/pbs-notify.sh and proven on the relay (PBS).

2026-09-11

  • Core Web Vitals of the four main sites — the Umami tag carries data-performance="true" on the hub, the landing page, this documentation and ISC Learn: LCP, INP, CLS, FCP and TTFB per page, from real visits (service). Heatmaps and replays left off, they would put 190 KB of rrweb on the pages being measured.
  • Saved Umami reports and a composed dashboard — funnel, goals, attribution and UTM on the landing page, journey and retention on the hub, retention on Learn; stats.isc-vs.ch/dashboard shows the four main sites on one screen. Both written by script (provisioning/stats/reports.py, dashboard.py), service.

2026-09-10

  • ISC Learn counts with Umami, Google Analytics removed — tag in Additional HTML → head, the theme's ganaidga4 emptied, caches purged, no downtime (setting, service).
  • Audience statistics for the ISC web sites — a self-hosted Umami on srv-stats (VM 121 on rumba, stats.isc-vs.ch), tracker tag deployed on all nine sites, campaign short links isc.hevs.ch/go/<slug> in the hub (service, site list). srv-freeipa (VM 106, .173) found undocumented on the node while allocating the address.

2026-09-09

  • Social-preview cards on the light tone — the picture Teams and Slack show for docs.isc-vs.ch and isc.hevs.ch was dark; both are regenerated light, and the hub's card finally prints isc.hevs.ch instead of the old GitHub Pages address. New filenames (og-card-light.png, og-image-light.png) because Teams caches the preview per URL; recipe in provisioning/web/og-cards/.
  • ISC Learn frontpage: ambient background, 44 px rhythm, menu with « ISC Hub » at root — visual pass with pmudry, six deploys (log).
  • ISC Learn: theme AJAX blocks requested at HTML parse time — frontpage tabs, blog, « Mes cours » and the course sidebar panel are fetched by two inline scripts instead of waiting for first.js; 32 s → 4.5 s under mobile throttling; the theme-file patch for the course panel is retired (log, how it works). The syntax highlighter filter serves its highlight.js files locally instead of from cdnjs (setting).
  • ISC Learn: 2026-27 students registered — 33 accounts and 371 per-module enrolments from the secretariat's workbook, by the new scripts; per-student enrolments replace the year cohorts (log, procedure).
  • Learn toggle sidebar drawn right at first paint, frontpage typewriter script repaired — both CSS/JS-side fixes on hannibal, no downtime (log).
  • Docker Hub pull-through registry for the k8s clustersrv-registry (CT 120 on rumba, registry.isc3:5000, 200 GB cache not backed up); the three k3s servers pull docker.io images through it, new VMs get the mirror config from cloud-init. Asked for by the Kubernetes course block starting 2026-09-17 (page).

2026-09-08

  • Roster groups renamed by kind, operators trimmed to threevpn-carnaval / vpn-rack-operators / vpn-rack-mgmt are routes, role-rack-admins / role-pve-auditor are rights; vpn-rack-operators now holds pmudry, yacine.said, adrien.reynard. On NetBird the k8s range carnaval-infra moved into net-carnaval-guests, so every vpn-carnaval member reaches the API and Headlamp; the hosts policy folded into operators-access. Peer login expiration was dropped to 1 h for the switch and put back to 48 h (access model, roster, your access, log).
  • Keycloak group calypso-users deleted — last remnant of the Calypso model, no consumer since the 2026-08-28 rebuild.
  • NetBird updated to 0.78.1 (dashboard v2.92.0, agent 0.78.1) over snapshot pre-0-78-1 (log).
  • carnaval8's NVMe hung, GPU lab moved to carnaval9 — the second Swissbit card suspended its pool on 2026-09-03 and froze gpu8 for five days; recovered by cold power cycle (pool ONLINE, scrub 0 errors), the lab rebuilt as gpu9 (VM 1169) and carnaval8 powered off pending the decision on its card (incident, log, todo).
  • A lab VM that changes node now gets that node's name — repointing an alias hands everyone a changed SSH host key; standing rule on the labs runbook.
  • Node-to-node SSH on carnaval explained — the near-empty /etc/pve/priv/known_hosts is a pre-8.2 vestige, not a broken cluster; PVE 9 uses a per-node file plus HostKeyAlias (access).
  • marcellus backed up by hand before decommissioning — SSH re-opened by Infomaniak (443 still closed); every web root, /etc (with its etckeeper history), logs, homes, the quarantine and InfluxDB archived, plus dumps of all MySQL, PostgreSQL 12 and tisc-db databases, 5.2 GB checksummed on pmudry's machine (C:\temp\marcellus_backup). Enough to rebuild a site elsewhere from its dump and uploads/; core, plugins and themes to be reinstalled fresh, not restored.

2026-09-07

  • isc-adm VM (104) brought into service — hosts the ISC registration manager on http://192.168.88.171:8000 over the VPN as a systemd unit, just deploy pulls and restarts; documented in the rumba guest table. Internal DNS name still to add (todo).
  • Yacine promoted to NetBird admin — a federated user lands as user whatever the roster says; the role is a third act after rack-admins and ADMINS=, now in the enrolment steps with provisioning/netbird/set-user-role.sh.
  • marcellus flagged by Infomaniak and cut off — Shadowserver sinkhole hit traced to a PyInstaller payload run by www-data, entry through the Duplicator installer left in the ikarus WordPress since 2021 and used on 2026-09-02; payload quarantined, ikarus disabled, host to be discarded (incident, containment log, plan in the todo).
  • hannibal hardened the same day — three accounts and their keys removed, yacine.said added as admin, sshd forwarding and root login off, Docker ports closed, 308 updates, automatic security updates back on; afternoon: ingegamez WordPress on auto-updates, Moodle code tree read-only for the web server. Kernel reboot pending (log, state on the server page).

2026-09-06

  • SSO admission gate generalised to every service — a login that will not be admitted is now refused in the realm with a sentence, on every client and not just NetBird; two rules, roster for machine access and isc affiliation for the published tools, with hevs dropped so the rest of the HEI counts as external there (exceptions go through roster.csv). Both oauth2-proxy gates moved to session_cookie_minimal after the new realm role pushed the session cookie past Stirling-PDF's 8 KB header limit (the gate, the user gate).
  • ISC Learn plugin updates — the six updates Moodle had been reporting, rehearsed on the mirror, 39 s on hannibal, no downtime (log).
  • New ISC Learn frontpage in production — ISC-sites visual language on the Moodle landing page, Custom CSS plus a builder page, nothing in the code (log, design).
  • ISC Learn content touch-ups — teaching-team page, « Infos » menu, phone pass, course images for 200.3, 304.2 and 305.2; scripts with --revert in provisioning/learn/ (log).
  • ISC Learn mirror refreshed — VM 107 on the day's dump and files; two fixes to the refresh script (log).

2026-09-05

  • ISC Learn performance pass — InnoDB buffer pool 128 MB → 2 GB, OPcache sized for the 5.2 tree, event log cut to one year, sessions and application cache moved to Redis; all online, no downtime (log).
  • ISC Learn upgraded to Moodle 5.2.2+ — rehearsed on the VM 107 mirror, done on hannibal at 23:15 with 118 s of downtime (log).
  • ISC Learn mirror refreshed, rumba copy versioned — VM 107 on the 2026-09-04 dump after a 37 GB delta pull; the copy moved to its own dataset hdd/hannibal-mirror, ZFS-snapshotted after every pull (log).
  • Frozen Aug-2024 Moodle removed from marcellus — 81.5 GB of files and the disabled vhost, no archive (every course is on hannibal); /srv 96 % → 9 %, databases left for a later drop (details).
  • Moodle course backups purged on hannibal — 800 of 1 194 automated backups removed, retention 3 / 1 by course age; ~92 GB back once the pre-purge snapshot goes (log). Closes the "shrink the Moodle course backups" item.

2026-09-03

  • gpu8 on carnaval8 — node powered on over Redfish (it already held templates 9008/9108 from the 2026-08-28 restore), VM 1168 stamped from the CUDA template with the A2 attached, names gpu8 / lab-1168; whole roster aboard, nvidia-smi sees the card. carnaval-guests.sh now publishes to carnaval810 too (labs runbook).
  • Yacine's Proxmox access checkedAdministrator on both clusters was already in place; the node Shell asking for a login: is PVE's root@pam-only rule, now noted on the rumba page.

2026-09-01

  • Yacine added as administrator — roster line staff;rack-admins;appliances-users;mgmt-users;carnaval-users, UID 1105 with a NAS home, root key on rumba, the carnaval cluster and srv-docker01, PVE user on both clusters, write access on ISC-HEI/dc-isc (enrolment, root SSH). The same run restored Adrien's root key on carnaval, lost in the 2026-08-28 rebuild.
  • gpu-01 specced — the EPYC 7443 listing of the G242-Z11, 3 × RTX PRO 4500, 8 × 32 GB RDIMM, ≈ €14 000; three dual-slot cards leave a Gen4 x16 for the 100 G NIC, which closes the slot-budget question (architecture §9).
  • A third R282-Z92 recorded as a cold spare — bought new in August 2026 outside ServerShop24: 2 × EPYC 7302 (Rome), 24 × U.2, no RAM, no drives. No spare CPU needed; its untested backplane is a possible fix for the two defective ones (Epyc, register, todo).

2026-08-30

  • VPN session window 24 h → 48 h — NetBird peer_login_expiration at 172800 s; the deprovisioning tail doubles with it (log).
  • CRS520 ordered (sw-100g); the ConnectX-4 Lx count covers every 25 G link, so the NIC line leaves the buy list. UPS candidate set: Eaton 9PX 8000i 3:1 on the 32 A way (architecture §7).

2026-08-29

  • The two R740xd are delivered — RG-789539: 2 × Gold 6244, 32 GB, BOSS 2 × 240 GB M.2 boot, 57416 rNDC each, plus 6 PM983, 36 carriers and 3 spare PSUs. Not racked, not yet assigned to pbs-01/filer-01 (register).
  • pve-03 accepted and installedepyc3, the R7515: its universal bays are cabled (12 U.2 slots enumerated), the chassis has four PCIe x16 slots and not two, iDRAC9 is Enterprise-licensed, and the 16 HPE SAS are 512-byte sector. 256 GB fitted from the epyc0 DIMMs; PVE installed from a USB key. The 300 GB SAS pair on the order sheet was never delivered, so rpool takes two of the 1.8 TB (todo, architecture §3).
  • 16 DIMMs harvested from epyc0 — it was at 1 TB in 32 slots; the 16 Micron modules in the Dimm0 slots go to pve-03 (8), epyc1 (6) and the shelf (2). epyc0 keeps 512 GiB, now at 3200 MHz (log).

2026-08-28

  • CIFS sudoers fix written, as a PR (audit F-CODE-1) — branch fix/cifs-sudoers-pin gives the per-user mount -t cifs rule in the Ansible common role a fixed nosuid,nodev,noexec,credentials=~/.smbcredentials option string instead of a free *. Not on main: the role only reaches the CALC machines, which Gregory owns.
  • Condensation alarm widened — reference raised from 16 to 19 °C (warn ≥ 17, crit ≥ 19): the dew point reached 17.3 °C without any condensation, and 16 fired on ~10 % of the hours (thresholds).
  • k3s teaching cluster on carnaval — three GPU VMs on carnaval02, API k8s.isc3, rebuilt by one script; carnaval-power.sh now re-quorates a partially powered fleet. Operation, what runs.
  • Calypso retired, every R630 is a carnaval node — the eight Calypso nodes and the three old carnaval nodes wiped and reinstalled as a fresh ten-voter PVE cluster (carnavalN = ex-calypsoN, address kept; carnaval5 held off, dead DIMM). Local homes rescued to rumba:/hdd/backup/calypso-preinstall/ first. Templates, CUDA image, SSO and status page redone; k3s is next. Operation, current state, install traps.
  • Users can no longer change their own email — a review of the day's changes found that a guests password could be turned into any future edu-ID identity: set the guest's address to a colleague's, and their first edu-ID login would be auto-linked into the guest account, groups and all. The address is admin-only now; nobody needed the write. Reasoning on Keycloak. The rest of the review held: guests carry guests and nothing else, no vpn-access, the VM is key-only SSH with two ports open.
  • Brute-force protection and a password policy on realm isc — off until now, which mattered little while every login was brokered to edu-ID and matters since the guests accounts authenticate with a password on a public endpoint. Temporary lockout only, 10 failures, 15-minute cap; policy is length-only because the passwords are generated, not chosen. Verified live: the counter increments and a legitimate login still passes. Settings and the clear-a-lockout command on Keycloak.
  • The realm records who changed what — admin events were off, so a change to a group membership left no trace; they are on now, with representations. Login events were kept 7 days, now 90. Asserted by provisioning/keycloak/realm-audit.sh; what is stored and where to read it on Keycloak. The realm's brute-force protection and password policy were still off at that point; see the entry above.
  • One login page for everything — the accounts without an edu-ID moved out of the user gate's htpasswd file and into realm isc, in the new group guests. That route authenticated before Keycloak, so it carried no password policy, no brute-force protection and no trace in the realm's logs, and it forced a second sign-in page in front of the themed one. Both gates now go straight to Keycloak. Managed with provisioning/keycloak/local-users.sh; details on Keycloak.

2026-08-27

  • srv-docker01 replaces the static-site container on test3.isc-vs.ch — VM 105 on rumba, Debian 13 + Docker, root by key for the rack administrators. The proxy forwards the public name to port 8080 of the VM and nothing else, behind the user gate for now. Page: srv-docker01.
  • Building 19 room surveyed — three three-phase sockets (1 × 32 A, 2 × 16 A), a distribution board in the room, and an existing condensate drain. The 32 A way reopens the UPS input choice that August's 2 × 16 A had closed, and E1/E2/E4 become existing rather than requested. The implantation plan and its printable PDF are at version 1.1.

2026-08-22

  • isc.hevs.ch serves the hub instead of redirecting to it — the address bar keeps the institutional name; /learn is unaffected (log, state).
  • The hub dropped the # from its URLs — history-mode routing, so pages are isc.hevs.ch/etudiant-es. Both hosts have the fallback that needs: FallbackResource on hannibal, and the build copies index.html to 404.html for GitHub Pages, which serves it with a 404 status but the right content. Links of the old /#/… form are rewritten client-side before the router reads the location, so they still land on the right page. The fleet links moved to isc.hevs.ch at the same time, making it the canonical of the two names.

2026-08-21

  • The root of isc.hevs.ch now 301s to the hub, which answers on its own name hub.isc-vs.ch; /learn is untouched. Closes the hub-at-the-root plan (log).
  • Three utility tools for students and staff — Stirling-PDF, IT-Tools and CyberChef on CT 117–119, the first user-facing wave off the service ideas list (service page). All three live on their public names. Only the PDF tool takes a login: the other two compute nothing server-side.
  • A second oauth2-proxy instance, the user gateisc or hevs, both claim-derived, so it needs no roster. The unit is now a systemd template and the admin gate runs as oauth2-proxy@admin from the same files.
  • The break-glass WireGuard peer drops 40–95 % of its return packets — reported as slow web browsing on an admin laptop; the cause was DNS going into the tunnel over a link losing half its packets, in dead windows of about 14 s (incident). Fixed for the reporter by moving to NetBird; the peer itself is still broken, open in todo → VPN and identity.

2026-08-20

  • Scenario 3 is retained, the two R740xd are ordered, and calypsomaster takes mgmt-01 — the scenario page is promoted to Architecture & execution (execution plan now in French), with a new executive summary on top and scenarios 1, 2, 4 archived under considered scenarios. The watcher role moved from rumba to calypsomaster for its iDRAC9 (why); rumba goes to lab/provisioning duty once emptied. Closes the off-site-PBS chassis decision in todo → backups.

2026-08-19

  • ikarus.snowmon.ch fixed — the alias was served over TLS without being on the marcellus lineage, so every visit by that name failed verification. Re-issued with it as a 15th SAN (dry-run first), valid to 2026-11-17; it now 301s to the canonical name. Turned up that certonly does not reload Apache despite installer = apache (details).
  • The two worst wiki imports are rewritten — the ISC Learn DRP is now a decision table of failure mode → restore path → duration, folding in what building the standing mirror taught (the delta-pull staging copy, and the mail muzzle that has to be removed deliberately on a cutover). ISC Inventory is 234 → 59 lines and its 3.2 MB of Snipe-IT screenshots are deleted; the UI's own labels carry the procedure.
  • Docs clean-up, third pass: the todo — 753 → 654 lines with nothing dropped; the sections grouped by provenance are dissolved into subject ones, two new ones (fabric, DRP) hold what the duplicates asked for, and the reference data moved to the pages that own it (Calypso).
  • Docs clean-up, second pass: structural noise — Snipe-IT moved to Services (on-site), drp.md folded into the ISC Learn plan, both with redirects; journal entries recording sidebar churn, dead wiki template stubs and five closed CALC items are gone.
  • Audit finding F-SEC-3 closed: no repo file is published any more — a Markdown link to a file outside the docs tree still emits it as a build asset, and three were live, including provisioning/uid/uid-map.csv with 136 rows of uid,gid,name,email. build/assets/files/ is now empty on a clean build; the rule is in CLAUDE.md.

2026-08-18

  • Docs clean-up, first pass: the false facts — 26 files. Six pages offered the legacy WireGuard as a live access path, including the UPS discharge runbook. Also: rumba's model in the RACK export, the node and R630 counts, NAS Volume 2 and the home count, and the "RouterOS 6 EOL" case against the CRS326 in three places.
  • The sso.isc-vs.ch break-glass no longer admits the LAN /24 — the matcher is NetBird-only (100.65.0.0/16) now, closing the audit finding that the legacy-WireGuard src-NAT let every legacy peer reach the Keycloak console ungated; the stale 100.79 range was corrected in the same edit. Deployed and verified. What changed for recovery: break-glass.
  • Both MikroTiks' armed sniffer filters cleared (CCR vpn_calypso, CRS326 ether1014) — the last piece of the audit's leftovers finding; the wifx.ch L2TP client had already gone 2026-08-07. Verified through Oxidized's same-hour pulls.
  • NetBird admin console gated — the vpn.isc-vs.ch dashboard now sits behind the admin gate; the agents' API, gRPC, relay and the embedded IdP stay open, break-glass mirrors the sso vhost. Verified from outside on all four paths. The gate's sign-in stays on sso.isc-vs.ch, so no proxy_prefix change was needed.
  • PBS backups verified after two weeks — every 03:00 vzdump OK (two jobs a night), 284 backup indexes, dedup factor ~38× (7.05 TB logical in 185 GB), daily GC green. hdd re-checked: 59 %.
  • Three audit code findings fixed in the repofollowups.exp now feeds passwords to sudo -S's prompt instead of remote argv (F-CODE-2; the two passwords still want rotating), the netdata kickstart.sh in both post-install scripts is pinned and sha256-verified (F-CODE-4), and NB_SETUP_PAT_ENABLED is off (F-CODE-5).
  • The hardware-purchase items are settled — the five PERC battery kits are on hand (fitting them is what remains, R630 hardware); the S+RJ10 cold spare and the pbs-01 SFP28 DACs were dropped by decision.
  • The three dead marcellus vhosts are disabled (gptisc, influxdb2, portainer — via a2dissite, configs kept). Checked before touching influxdb2: no legitimate consumer, only scanner noise (inventory); the InfluxDB service and data stay in place. Live sites verified unaffected.
  • Hasdrubal settled: destroyed — NXDOMAIN on both zones and absent from the August Infomaniak inventory. The two references in change.md and drp/isclearn.md now carry what mattered (the Infomaniak system snapshot is unreliable) without naming a machine that no longer exists.
  • A second admin address added to the relay allowlist — applied on CT 112 and verified by SMTP dialogue (250 for it, 554 for an outside address); every relay consumer inherits it. The list lives in RECIPIENTS in provisioning/mail/deploy-relay.sh.
  • provisioning/web/ and provisioning/status/ got their READMEs — the poller's frozen CSV column order and the hand-pushed Telegram token file are now written down (IaC item 7); a Carnaval row was added to the machine table.

2026-08-17

  • Acquisition register created from the 13 ServerShop24 invoices (2023–2026) — provenance and component tallies per invoice, and the R630 count reconciled at eighteen: purchases.

2026-08-16

  • .isc3 alias domain added — every .calypso name also resolves as .isc3 (41 twin records on the CCR2004, NetBird nameserver group matching both); the 72 legacy WireGuard peers' client-dns and the deploy scripts' --searchdomain fixed in the same pass. Details: history.
  • Retired admins / teachers NetBird groups removed — empty jwt husks left by the identity collapse, referenced by no policy; the sessions that used to re-create them are long gone.
  • NetBird updated to 0.77.0 — server, dashboard (v2.91.1) and the routing peer's agent; store migrated clean, access model unchanged, a few seconds of interruption per restart. Details and the new detached update path: history.
  • Scheduled jobs wired to Healthchecks — six checks (both vzdump runs, PBS sync + verify, the hdd scrub via zed, Oxidized liveness), hooks installed by provisioning/healthchecks/wire-jobs.sh; mechanisms in the checks table. NAS-side tasks remain (todo).
  • Healthchecks + ntfy deployed — CT 115/116, closing the admin wave (4/4). Healthchecks alerts by email and ntfy (both paths verified); wiring the actual jobs is on the todo. Details: healthchecks, ntfy.
  • NetBox seeded — rack, 30 devices with their addresses, 7 prefixes and 13 VMs/CTs imported from the docs by provisioning/netbox/populate.py; details on the service page, gaps (cables, serials, Ansible) on the todo.
  • Oxidized deployed — CT 114 srv-oxidized, both MikroTiks pulled hourly into a local git repo as a dedicated read-only account; admin wave 2/4. Details on the service page; the FS S3600 gap is on the todo.
  • NetBox deployed — CT 113 srv-netbox, internal-only, Keycloak SSO gated to rack-admins; first of the admin wave (NetBox → Oxidized → Healthchecks + ntfy) picked from the service ideas. Details on the service page; populating it is on the todo.
  • Vaultwarden retired — CT 111 destroyed, vhost, Keycloak client and DNS records removed; it never held a shared secret, the secrets plan being SOPS + age only since 2026-08-13 (rumba history).
  • Condensation alarm on srv-status — the poller now publishes the dew point and alerts when it nears the chilled-water inlet (16–17 °C measured; warn ≥ 14 °C, crit ≥ 16 °C), on the humidity tile and over Telegram (thresholds). The past week peaked at 17.3 °C dew point — above the inlet.
  • Scenario 3: homes filer decided, backup down to one box — the student homes go to filer-01, a dedicated standalone filer (twin R740xd of the PBS chassis, closing the §4 A/B question), and the pull-synced second PBS box is dropped for a single hardened pbs-01 in 307; scenario 4's shelf quantities halved to match (§4, §5).
  • Carnaval lab VMs became cohort machines — accounts and root now come from the roster, keys from an ISC³-only file, and every VM gets DNS names (lab-<vmid> + weekly alias); gpu0gpu2 deployed, one per node with its GPU (stack page).
  • CUDA image pipeline: build once, publish via the NAS--cuda-prepare/--cuda-seal with a manual-tuning window, dump on the new nas-library share as the cold copy (pipeline, library); templates rebuilt with clang/cmake/nvtop. Trap fixed: the build VM is only reachable through the node, the old builder polled it from the operator's machine and concluded it never booted.

2026-08-15

  • rumba rebooted — the UPS USB link did not come back (history, incident).
  • NetBird was down ~40 min after the reboot — the no-hairpin hosts entry is now persisted in the cloud-init template (incident).
  • A system engineer joins ISC³ end August — "bus factor 1" is restated as a two-person team wherever it carried a decision, in all three scenarios; scenario 3's phase 0 gains the two-person inspection sprint that turns the buy list into a number (scenarios).
  • The R7515 order and the NIC shelf folded into all three scenarios — hybrid backplane confirmed as 12 SAS + 12 universal, two 300 GB SAS shipped so a SAS rpool costs nothing, RAM corrected from CHF 400–550 to ≈ €2 350; 6 ConnectX-4 Lx (25 G ceiling) plus ≥ 6 spare ConnectX-6 mean no 100 G card is bought in any scenario. Scenario 3's pve-03 takes all 12 PM983 for nvme_pool (10.5 TiB) with bulk down to 10 disks, base subtotal ≈ €5 750 (scenarios).
  • The ConnectX-6 are single-port, which breaks scenario 1 §6.2epyc1's BMC gives MT28908/MT4123 ConnectX-6 VPI, one QSFP56 port, Gen4 x16, fw 20.31.23.54, and the spares are the same model. The "port 2 goes back-to-back to the other node" replication link has no port 2; it now takes a second spare card per EPYC node, so still nothing bought but one more Gen4 x16 slot each (§6.2). Scenario 3 is unaffected — one port per node into the CRS520 is all it asks for.
  • Scenario 2's chassis is not in stock — no Dell 12 × LFF listed at ServerShop24 on 14 August, nor any MD1200/MD1400 shelf, and the 8 TB SATA line it prices is gone as well; meanwhile the 24-bay all-NVMe SKU its §3 dismissed as uncommon is stocked at €1 639 (option 2).

2026-08-14

  • Adrien Reynard has root SSH on rumba — his key added to /etc/pve/priv/authorized_keys. Who holds root on a PVE node is now the git-tracked list in provisioning/pve/root-ssh-keys.sh (rumba, access register).
  • Read-only Proxmox is a roster entitlement nowpve-auditors grants PVEAuditor on / through the PVE group pve-auditors-isc, the first rights group managed two-way from the roster (SSO on the node, resource groups). Asserted on both clusters — rumba, and carnaval from carnaval0, which covers all three nodes.
  • Adrien Reynard's access correctedappliances-users (no reach to the 88 subnet, so rumba was unreachable over the VPN) plus pve-auditors, both from the roster. He had been put in rack-admins, which grants no reach at all and would have made him Administrator on rumba at his first login; taken out of it in the console, as that group is add-only in the sync.

2026-08-13

  • calypsomaster and rumba inventoried over SSH and Redfish — both were mis-modelled in the docs. calypsomaster is a Dell R740 with 8 SFF SAS/SATA bays, not the R740XD §5.2 assumed, so it cannot hold the 10.5 TiB second backup copy (calypso, todo). rumba is a Precision 7920 Rack, has every DIMM slot populated, five free PCIe slots and an uncabled ConnectX-6 nobody had recorded (rumba).
  • Storage & backup option 2 written — two identical 12-bay LFF chassis for pbs-02 and a production filer, with the R7515-into-production variant and its phase-4 deadline (option 2). Nothing ordered.
  • The storage plan for the whole estate is settled — dropping the 19 TB tank turns the dead backplanes into an inventory problem; the allocation across every machine is in §5.2. Nothing fitted yet (todo).
  • epyc1 checked over end to end — CPU, storage, PCIe, thermals and firmware all clean (nvme_pool at 4.7/7.8 GB/s, no ECC events across 167 GiB written). The missing 32 GB turns out to be a firmware map-out of socket P1's channel P, and enabling soft PPR did not recover it, leaving a CMOS clear and a CPU swap. An unused ConnectX-6 and an unplugged PSU1 also turned up (checkup, Epyc).
  • The secrets plan drops Vaultwarden — one scheme for everything, SOPS + age over secretzone/, with a break-glass age key in the sealed offline set; the deployed instance (CT 111) becomes a removal item (secrets management, todo). Nothing migrated yet.
  • The hub keeps its host out of its name — the isc.hevs.ch root redirect will point at hub.isc-vs.ch served by GitHub Pages, not by srv-web01 as first recorded: nothing to deploy on the rack, and a later move to another host costs one CNAME instead of edits on hannibal and across the fleet (plan). Still nothing executed.

2026-08-12

  • epyc1 documented and its storage built — a Proxmox VE standalone node on 192.168.88.46, BMC pinned to 192.168.88.12 (epyc1-bmc), with its three 3.2 TB NVMe as nvme_pool: RAIDZ1, 5.68 TiB usable, from provisioning/pve/epyc1-nvme-pool.sh. New page: Epyc. epyc0 is in repair. Credentials in secretzone/epyc.md; root password rotation and the missing mail path are in the todo.
  • The racked Tango node answers on 192.168.88.43, not the documented .20 — its MAC has no reservation (Tango, decision in the todo).
  • Both EPYC U.2 backplanes are defective, different bays on each, so the PM983 drives stay out and the 19 TB tank target is blocked. Diagnosis, the power-only LED test that localises it, and the bay/CPLD reference are on the Epyc page; the four options, cheapest first, are in the todo. Also recorded there: epyc1's dead memory channel and the deliberate empty DIMM slots that look like faults but are not.

2026-08-11

  • SInf also opened TCP 587 to Infomaniak — verified from rumba; 587 to any other host and 25 to anything still time out, so it adds a port, not a second provider. The relay stays on 465/implicit TLS (Email, uplink restrictions).
  • Wording pass over the whole docs tree — flattened PR-voice emphasis in 75 files (bold used to insist, intensifiers, verdict sentences) without touching facts, headings, numbers or links. The rule now sits in CLAUDE.md.
  • Factual contradictions found by that pass, corrected — node counts (the rack has 8 Calypso and 3 Carnaval, neither page had it right), the NAS exports being async not sync, the frozen Moodle's size, NetBird's session window, the CRS326 having been upgraded after all, the WireGuard peer count (74, from the 2026-08-10 config export), two ::: closers that swallowed the paragraph after them, and assorted malformed markup. What still needs a decision is in the todo.
  • UPS replugs failed, NUT still blind — both connectors reseated, neither produced a kernel event while port 13 reports healthy, leaving the cable, the UPS's own port, or a dead port 13 (incident). The unit does have a DB9, so repair-or-move-to-serial is now a real choice rather than a hypothetical.
  • NetBird deprovisioning documented properly — a NetBird-side block cuts immediately, disabling the upstream account alone leaves a 24 h tail, and the way to revoke now is to end the session by hand in Keycloak (netbird).
  • Decision: the root of isc.hevs.ch will point at the hub (hub.isc-vs.ch), /learn untouched — recorded in the learn migration plan, work items on the todo; nothing executed yet.

2026-08-10

  • tMatch is gone from marcellus — everything behind it deleted on request without a backup, which also closed 16 root-equivalent entry points and the FreeIPA-on-0.0.0.0 audit finding. The source is safe on GitHub; the database is not (details).
  • Marcellus's shared certificate too — re-issued without the NXDOMAIN apt.chezmoicamarche.ch, 18 names valid to 2026-11-08, all verified from outside; its renewal was due around 2026-09-04 and would have failed all 19 vhosts (marcellus). Both shared lineages are now clean.
  • hannibal's shared certificate is fixed before it could take seven vhosts down — re-issued without wiki.isc-vs.ch, the missing Apache reload hook added, standing rule written (log, rule).
  • The rack can send email, for the first time — SInf opened 465 to Infomaniak and everything goes through CT 112 srv-mail, the Postfix relay that holds the only copy of the credential. Six consumers wired the same day, zfs-zed and smartd among them, which had been mailing into a void (Email, rumba log).
  • secretzone/ moved to the repo root (with secretzone/calc/ for CALC), out of both Docusaurus trees — the published site can no longer expose it even by misconfiguration; all referencing docs and provisioning scripts updated (secrets management).

2026-08-09

  • The UPS's USB link wedged for 7 h, blinding NUT from 05:03; a usbreset brought both interfaces back with the UPS itself never at risk. NUT's insufficient permissions on everything means a wedged link, not a permissions problem (recovery, incident).
  • The CRS326 is off RouterOS 6 — 6.49.19 → 6.49.20 → 7.21.5 → 7.23.3 in one afternoon, RouterBOOT matched, so both MikroTiks now run the same current firmware; config conversion intact (details). Two standing rules came out of it: a MikroTik upgrade cannot be predicted from npk size versus free space, and v7 is offered to a v6 device only on the testing channel (both). Still open: L3 offload, now within reach (todo).
  • The CRS326 can reach the world again, after two independent breakages: its DNS pointed at the retired SInf gateway 172.30.7.1, and its own input chain ended in drop with no established,related accept, so replies to traffic it originated were dropped by itself — the forward chain does have that rule, which is why everything behind the switch was online while the switch was not. NTP syncs and check-for-updates works; the rule is on the network page.
  • The routers have a remote rescue path, and a smaller management surface — RoMON enabled on both with a shared secret and forbidden on the WAN port, so WinBox reaches the CRS326 through the CCR when its IP config is broken (exercised: discovery finds it at one hop). telnet, ftp and plain www disabled on both — part of F-NET-4; api (8728) stays, routeros-api.py needs it (todo, audit notes in secretzone/security-audit-2026-08.md).
  • CCR2004 upgraded to RouterOS 7.23.3 (from 7.20.4) with its RouterBOOT to match — two reboots, under a minute each, config and services verified end to end (details).
  • The CCR's pre-WireGuard remote access is deleted — L2TP server (use-ipsec + shared secret), the vpn PPP account and the wifx.ch L2TP client to the former ISP. All four were already disabled and never used; they only survived as three cleartext passwords in the config (network).
  • MikroTik backups are now complete — both devices re-exported with sensitive values (Wireguard private keys, the wifx.ch uplink password, the L2TP/IPsec secret) and their binary backups copied off for the first time, so user accounts and certificates are covered too. One script does all four steps: provisioning/network/mikrotik-backup.sh (process).

2026-08-07

  • ISC Learn migration plan written — target: Infomaniak Serveur Cloud managé (proposal); prod facts refreshed read-only (log).
  • hannibal's fail2ban now bans — the sshd jail watched port 2002 while sshd listens on 20002; fixed, and ubuntu given a console break-glass password (log).
  • Marcellus is inventoried — read-only survey of the legacy VPS: 19 vhosts (5 answering 503), 6 containers, ChirpStack, four PHP-FPM versions and a frozen Aug-2024 Moodle using ~81 GB of a 96 %-full /srv (Marcellus). Turned up a dated [HIGH]: a name on its shared 19-vhost certificate no longer resolves, which would fail the renewal for all of them around 2026-09-04. Decommissioning otherwise blocks on naming the owners of the non-ISC vhosts (todo).
  • The pagode pair is renamed tango — the name fits a two-machine pair, and the old name followed no convention. The planned CALC@HEI login node that already held the name Tango becomes mambo (secretzone/calc/mambo.md). Also placed at U36 front on the schematic, where it actually belongs (Tango).
  • What a Proxmox cluster brings is now written down — the docs only carried the costs (quorum, join/leave constraints, HA vs. thermal shutdown); new section splitting what a plain cluster gives from what needs shared storage (§4.1), plus cluster/corosync/quorum/ QDevice in the glossary — renamed from "Acronym glossary" to hold terms too. Stale PBS retention (7/4/3 → 14/8/6) and datastore size fixed on the backups and rumba pages.
  • carnaval1's NVMe controller hung and took fast-vm down for 17 h — recovered by a cold power cycle from the iDRAC; the pool scrubbed clean and the card is back in service, on watch (incident, recovery sequence).
  • The legacy MikroTik WireGuard fleet is off — 12 peers removed and every other one disabled, leaving wg13 as the only enabled peer. Attribution table, what was removed and the break-glass consequence: secretzone/wireguard.md.
  • Security audit of the whole architecture — read-only; findings and the ordered remediation list in secretzone/security-audit-2026-08.md, open items marked by severity on the todo. Also stopped two files linked from the calc pages being published as site assets.

2026-08-06

  • The disposable GitLab test pair is destroyed — CT 105 srv-gitlab (with its secretzone copy and full-rights API token) and VM 106 srv-runner01 are gone, along with their five DNS entries and the teardown checklist page itself; what the test proved and the deltas for a real instance: history entry.
  • Two group deletions silently took five VPN policies with them — ~1 h of access-model outage, restored in one replay-config.py run from the committed snapshot (its first real recovery); the cascade is now a standing rule on the service page (incident).
  • The identity model collapsed to students / staff — identity now grants admission only; the old teachers / admins reach became the entitlements appliances-users (appliance subnet
    • carnaval PVE hosts) and mgmt-users (iDRACs), so every grant is per-resource. NetBird policies repointed by provisioning/netbird/migrate-role-groups.py (access model, role groups). First staff member: Marta Rende.
  • Enrolling is now a documented two-step — roster line + sync, with the un-enroll path and its deprovisioning tail: enrolling people.
  • The ISC wiki is retired — DokuWiki archived off hannibal (rumba:/hdd/backup/wiki-archive/, where it is) and wiki.isc-vs.ch turned into a 301 vhost on srv-web01 that maps the old doku.php?id=… links onto their new pages (the vhost).
  • The documentation has its own name: docs.isc-vs.ch — GitHub Pages custom domain (static/CNAME), so the site is served at the root instead of under /dc-isc/. GitHub redirects isc-hei.github.io/dc-isc/<path> there by itself, so old links keep working.
  • A non-enrolled edu-ID login to the VPN now gets told why — Keycloak denies NetBird logins lacking the roster-derived vpn-access role with a themed EN/FR "not enrolled — contact the staff" page, instead of NetBird's bare error after a successful login (the VPN gate, provisioning/keycloak/vpn-access-gate.sh). The first binding broke every edu-ID login for ~20 min — the trap and its fix are a standing rule on the Keycloak page.
  • Pre-created roster accounts now link silently on the first edu-ID login — the eduid-autolink first-broker flow replaces a verification Keycloak could not deliver, and edu-ID becomes authoritative for the stored profile. Validated end to end with two students, VPN gate included (details).

2026-08-05

  • DHCP is off on the node subnet — the CRS326 pool overlapped the carnaval lab-VM range while the VMs are assigned statically, and nothing had used the server since the MAAS client that held its only lease disappeared. Pool moved to 192.168.91.200-.240 so a careless re-enable cannot land in .128.191, both servers disabled (todo → network follow-ups).
  • Lab VMs get their address and their name by themselvescarnaval-lab-vm.sh allocates the lowest free octet in 192.168.91.128.191 from the guests the cluster actually runs, registers lab-<vmid>.calypso on the CCR2004, and --destroy drops both. The range is the NetBird resource, so a VM can no longer land where no VPN client reaches it (how to run it).
  • Numeric UIDs are issued from a register, and Ansible reads itprovisioning/uid/uid-map.csv seeded from the ownership NFS actually compares, with uid-alloc.py to allocate and verify and create-home.sh to make a home with the number it holds. Applying it will renumber 13 rack accounts, uid-alloc.py --check-ansible lists them (UIDs).
  • elia.pacioni renumbered 1027 → 1100 — 1027 is IscAdmin in the NAS's own /etc/passwd, so DSM rendered that student home as owned by the admin account. Moved on the NAS and the three nodes together, so nothing drifts apart.
  • Seven departed accounts removed from the Calypso nodes, their SSH keys and the register, and the last two /home trees dropped from the carnaval pre-install rescue on rumba. Their numbers are free again; nothing else in the 203 rescued homes was touched.
  • The roster now drives VPN access per resource, not per role. roster.csv became one line per person (email,groups[,nasname]), 18 third-years and 3 teachers are in, and the flat 192.168.91.0/24 grant was split so calypso-users and carnaval-users mean different things (role groups, access model).
  • UID bands widened and given a hole for DSM — the old plan ran out after 500 students and collided with the Synology's own accounts (UID allocation).
  • The SSO login page now names the service you are signing in to and hides the local username/password form behind a disclosure — edu-ID is the only thing the card offers at first glance, which matches reality: realm isc holds no password at all. Theme in provisioning/keycloak/theme/ (Keycloak page).
  • New index: service access paths — the URL or address of every service and what it asks you for, in one table set, linking to the page that owns each one.
  • Proxmox now takes an edu-ID login through Keycloak, on rumba and on the carnaval cluster — realm isc (client proxmox), root@pam untouched and still preselected. Only pre-created users are admitted (--autocreate 0), and the sole Administrator grant on each cluster is the PVE group rack-admins-isc, filled from the Keycloak groups claim at every login — so revoking there revokes here. PVE appends the realm name to claim groups, one of the traps recorded on the node page.
  • The recurring netdata UPS mails were a false alarm, now silenced at the sourceupsd_ups_last_collected_secs flapped 3–4×/hour (40 pairs in two days) on stock's 5 s threshold, which the driver's stale-data spells cross routinely; it now warns at an absolute 60 s (why, and why the collector must not be reconfigured). The mail path was Netdata Cloud — rumba is claimed, and its own postfix has delivered nothing in days.
  • Carnaval nodes added to the rack status page, with a netdata deep-link — the poller was still calling them calypso810; any machine card flagged netdata: true now links straight to its dashboard (how).
  • No lab VM mounts the whole homes tree any more — the pre-2026-08-05 VMs are gone (the cluster holds only the six templates), so that item closes; the same exposure on the Calypso nodes, where docker makes students root-equivalent, stays open until the rebuild.
  • --sudo grants root, which it did not beforeuseradd leaves the account with no password at all, so group membership alone left sudo prompting for a credential that cannot exist. carnaval-lab-vm.sh writes a NOPASSWD drop-in per granted name, the way cloud-init already does for ubuntu (how to run it).
  • A GPU guest reads as 100 % memory used, and that is expected — VFIO pins every page and passthrough forces balloon: 0, so PVE reports the qemu process RSS rather than guest usage: 62 GB against 1.3 GB actually in use (why). Same pass corrected two stale claims — ubuntu-2404-cuda exists on all three nodes, and the reference guest VM 110 is gone, so all three GPUs are free.
  • Lab-VM UIDs now come from the NAS, not a login nodecarnaval-lab-vm.sh stats the home directory, because the ownership NFS enforces is the only source that cannot be stale: a Calypso node's passwd disagreed for 4 of 67 accounts, and a wrong number fails silently. It also refuses a home owned by uid 0. The permanent fix — issuing UIDs once instead of rediscovering them — is designed in UID allocation; the four drifted accounts and the unused groupe1groupe4 dirs are open items.
  • --sudo takes names--sudo=pmudry gives root to one person on a lab VM instead of the whole class, so a teacher can hold a privileged account of their own with their real UID and NAS home (how to run it).
  • Two home-ownership repairs on the FS2500olivier.amacker's home was owned root:root, restored to his 10011:10011; the four unused groupe1groupe4 project directories and their groups on calypso0/1/2 are gone (open items).
  • A NetBird restart destroyed its own storeencryptionKey was empty, so the server had been minting a throwaway key at every start; ~90 min VPN outage, carried by the legacy WireGuard, and the store rebuilt from config-snapshot.json with the new replay-config.py (incident, what the account looks like now).
  • The VPN logs in through Keycloak, and access follows the rosternetbird up goes to ISC SSO then edu-ID, and admission is membership of a role group instead of a manual approval nobody could see (how it works, what it took). Remaining: open items.
  • edu-ID logout no longer ends on a French error page — the provider's logout URL is now empty, so logout ends at Keycloak and returns to the client (why).
  • NetBird lazy connections off — they were stranding a connected peer with correct routes; relay-only with three peers gains nothing from them (detail).
  • History sub-pages made uniform — one name for all of them: <page>-history.md, titled "history & operations", newest entry first. nas-operations.md was renamed and the rumba page's sections flipped to newest-first; the convention is now in the document-work skill.
  • Ansible lives in this repositoryprovisioning/ansible/, playbooks and their configuration together, laid out the way the setup script used to assemble them so no playbook changed. The two GitHub repos are archived. Why: the roster and SSH keys sat in one repo while everything consuming them sat elsewhere (how it works now).
  • Lab VMs no longer expose the whole cohort's homes — one NFSv4.1 subdirectory mount per student instead of the homes tree, and sudo is now opt-in (--sudo), since root on any box in the node subnet can impersonate any UID over sec=sys NFS. Verified: a student reads their own home, another student's path does not exist (how to run it).
  • The MikroTiks are driven over their API, not SSHprovisioning/network/routeros-api.py. RouterOS echoes log lines into interactive sessions (stock action=echo topics=critical), which races every scripted pattern-match against the ~10 s login timeout; the API has no terminal and answers in under a second. Used it to disable the Dude server, the source of the constant login failure … via winbox entries (it polls with credentials that no longer match).
  • Student-home NFS mounts work again — the 2026-08-02 hardening had made every new NFSv4.1 mount of calypso_homes/homes fail while the live ones kept working, so a node reboot would have lost its homes. Fixed with one non-inheriting r-x ACL entry for guest on the share root, keeping root_squash: what it was and why that fix.
  • NAS writes went from 65 to 102–109 MB/s — both exports switched to async, which was the bottleneck; the client-side sync everyone assumed was the problem costs nothing on top (measurements and the durability trade-off). Writes now match reads at 1 GbE wire speed. Also corrected on the way: the claim that the FS2500 refuses NFSv4 was wrong.

2026-08-04

  • Carnaval has an operator runbookRunning labs on carnaval: handing a VM to a student, attaching a GPU, cohort loops, end-of-semester cleanup, and rebuilding the cluster from the repo. Also on 2026-08-04, the duplicate dhcp_test server on the CRS326 was disabled, so a VM can no longer be handed a node's address.
  • Carnaval can hand out lab VMs, GPU included — base and CUDA templates on the three nodes plus the cohort pools; clone to SSH in ~30 s, and a T4 passed through ran a CUDA kernel clean. Conventions, the VMID/address plan and the passthrough recipe: cluster page. Remaining: todo.
  • The playground cluster existscalypso810 wiped and rebuilt as the quorate 3-node Proxmox cluster carnaval, which starts phase 3; the old /home was rescued to rumba first. Two traps written up where they belong: the BIOS cannot boot the NVMe cards, and iDRAC8 needs its own install path.
  • PSU criticals cleared on calypso0calypso3, and the R630 fleet swept for PERC batteries — an iDRAC GracefulRestart re-inventories the chassis and drops a deliberately absent PSU (CriticalOK); clearing the SEL alone does not, and it never clears an absent BBU. Three nodes (calypso1, calypso2, carnaval0) are missing theirs and stay Critical until the kits arrive — table, order and traps in ops todo → R630 hardware.
  • Admin panels are behind a login instead of an IP range, and the split-horizon DNS layer is goneoauth2-proxy on srv-web01 guards the admin paths against the Keycloak group rack-admins, so they work from anywhere with no VPN and no client-side DNS trickery, and the layer that forced an internal answer for a public name is deleted. A LAN/VPN break-glass branch stays, since the gate authenticates against Keycloak itself (how it works).
  • Password vault live, and it is Keycloak's first consumer — Vaultwarden 1.37.1 as CT 111 (srv-vaultwarden), published at vault.isc-vs.ch, login through the realm isc (SSO_ONLY, invite-only). Keycloak login does not replace the per-user master password. (Retired 2026-08-16 without entering service.) Needed trustEmail=true on the eduid provider, otherwise brokered users arrive unverified and Vaultwarden refuses the signup.
  • Rumba's own configuration is now codeprovisioning/pve/post-install.sh covers everything the installer leaves out: deb822 repos, the hdd RAIDZ2 pool, the five storages, both nightly backup jobs, netdata. Idempotent, so it doubles as a check that the node still matches its page. Closes the first of the IaC audit gaps; items 4–7 there remain.
  • Keycloak realm and NetBird access model are snapshotted into gitexport-realm.shrealm-isc.json and export-config.shconfig-snapshot.json, secrets stripped, to be re-run after any console/dashboard change and committed. Both are diff tools, not restore paths: neither API can import what it exports — Keycloak, NetBird.
  • edu-ID federated login works and feeds three institution groups — the resource registry approved hes-so_isc3_vs_oidc_sso, a real login brokers a user, and hes-so / hevs / isc are filled from the claims (how). Marking an ISC member took an RR amendment; the measured claim values and the four traps are in secretzone/eduid-oidc.md.
  • SSO login page now wears the ISC brand — Keycloakify theme (petal colors, logo, dark mode) built in provisioning/keycloak/theme/, deployed with deploy-theme.sh, loginTheme=isc on the realm. Build/deploy notes on the Keycloak page; provider jars now survive in-place upgrades.
  • Rumba's guests now back up off-host — PBS 4.2 as a VMM VM on the FS2500, nightly job for all guests, chain verified with a live restore-capable backup; Volume 2 became PBS-only and the local job was trimmed to keep-daily=3 so hdd stops overflowing. The rationale and the VMM/token traps are on the service page.

2026-08-03

  • Netdata's stock UPS battery-charge alarm silenced — it keys on nutdrv_qx's unreliable charge estimate, and the UPS firmware intermittently zero-fills battery.voltage in its Q* replies, so the alarm fired with the UPS on line. Override now installed by deploy-ups.sh; details.
  • Rumba moved onto the rack UPS (both PSUs, ~5 s per swap, no reboot). UPS load 14 % ≈ 265 W — rumba is the only thing on it, and the UPS input is itself a PDU outlet. Measurements in the rumba history. Still open: NUT is monitor-only, so a long outage hard-kills rumba — todo.
  • UPS status on the status page — NUT polled every 30 s, UPS tile + load chart + history columns, Telegram alerts keyed on ups.status only (never on the estimated runtime).
  • UPS real battery telemetry decoded — the second USB interface (HID Power Device) reports a measured autonomy of 55 min at rumba's load; NUT's own charge/runtime estimates are unreliable on this unit. Read with ups-read-hid; details and traps in the HID section.
  • NAS given one dedicated 10 GbE port per subnet — second cable, eth3 static 192.168.91.250/24; rumba ↔ NAS measures 1.1 GB/s. Bonding rejected on purpose. NAS history; rebuilt nodes must follow how a node must mount it.
  • UPS monitored over USB (NUT) — the BlueWalker has no network interface, so USB from rumba is the only management path. Service page; runnables in provisioning/ups/.
  • Rumba's update backlog cleared — 0 pending, pve-manager 9.2.6 / kernel 7.0.14-8-pve.
  • Docs reorganized into present / past / future tiers — machine/service pages hold current state only; dated logs moved to history sub-pages; incidents in the incident log; open items consolidated in the todo.
  • NetBird masquerade turned off — appliances now see each user's own overlay address (100.79.x.x), so logs and per-IP protections act per user; the July 2026 lockout pattern is gone on this path. Service page.
  • Keycloak identity broker deployed — CT 110, https://sso.isc-vs.ch, realm isc, admin console LAN/VPN-only. Service page; runnables in provisioning/keycloak/.
  • SWITCH edu-ID OIDC client registeredhes-so_isc3_vs_oidc_sso, awaiting RRA approval. One client for the broker; GitLab, Grafana and Moodle become Keycloak clients with no further AAI paperwork. Record and constraints in the secretzone (eduid-oidc.md).
  • NetBird: edu-ID needs no licence — external OIDC login and JWT group sync are Community features; only SCIM directory sync is Enterprise. Do not buy it. Details.
  • Proxmox web UI given the ISC typography — self-hosted Manrope/Inter, idempotent deploy with --revert. Run check-coverage.sh after every pve-manager/extjs upgrade. Rumba history.
  • NetBird group profiles validated client-side — admin and student profiles match the access model exactly, split tunnel confirmed. NetBird history.
  • NetBird deployed in productionvpn.isc-vs.ch, VM 109, everything on the existing 443; relay-only is the permanent mode (the uplink drops inbound UDP 3478). The MikroTik WireGuard stays as the emergency/admin path. Service page.

2026-08-02

  • Network performance audit run and fixes applied — node↔node is wire-speed 1 GbE; the 9188 path was software-routed at 38.7 MB/s, fixed with on-link routes (61.5 MB/s NFS writes), a fasttrack-connection rule and the layer-3-and-4 bond hash policy. No new hardware needed. Network history; resulting config in the storage fast path.
  • Calypsomaster audited (read-only): nothing load-bearing runs on it. Findings and the conversion-to-second-PVE-node plan in the todo.
  • Ansible repos clarified: ansible-playbooks (+ private -conf) are the current fleet management; the copies on calypsomaster are stale duplicates of the archived repo.
  • FS2500 given its own NAS page with the first disk-level readout — Volume 2's Toshibas are 7.4 years old, older than the NAS itself.
  • FS2500 NFS exports hardened — three real holes closed (a /24 mask that admitted every VPN user as root, no_root_squash on the node rule, delete rights on the homes ACL). Details.
  • FS2500 data scrubbing run for the first time (Volume 1 had never been scrubbed in 1.8 years) — 0 errors, mismatch_cnt=0; quarterly schedule enabled on both pools. Details.
  • FS2500 Volume 2 rebuilt from 7-disk RAID 5 to 6-disk RAID 6 + a hot spare covering both pools — survives two failures, and Volume 1 gets automatic rebuilds for the first time. Capacity 11 → 6.7 TB, at no cost (the volume was empty). Details.
  • FS2500 rumba_backup stub removed — 6.7 TB of free RAID 6 available for the DS923↔rack replication goal (still in the todo).

Earlier in 2026

  • Patch panel for room 23N307 searched and bought (server-room list). Placement constraint to respect at install time: outside the rack, not inside — rack space will be premium.