Ops journal
Completed actions, newest first. This is where finished items from the todo land, so that page only ever contains live work. Each entry links to the page holding the details.
2026-09-21
- phpIPAM deployment -
srv-phpipamdeployed on VM 124 rumba, published athttp://srv-phpipam.isc3. TLS has not been implemented yet. Service discovery is performed by a cron job running every 30 minutes. Documentation is available here. - phpIPAM - mailing - Add phpIPAM ip in the networks list in mail system
2026-09-18
- TISC editor migrated off marcellus to its own guest —
srv-tisc-editor(LXC 123, rumba), published athttps://tisc.isc-vs.chthroughsrv-web01; DHCP reservation + internal DNS added on the CCR2004, public CNAME at Infomaniak,marcellus.isc-vs.chDNS removed (service, history). A baked-inAUTH_URLpointing at the guest's private IP made every login redirect unreachable from outside the VPN; fixed the same day. Marcellus decommission list: ops todo → hosting. - Slurm and MPI on the GPU VMs — one scheduler per VM, the card as
gres/gpu:1, OpenMPI over PMIx, in the CUDA recipe (Slurm). Template rebuilt and published to the five online nodes,k8s0–2destroyed on the way,gpu0–2andgpu9re-stamped (history); the five offline nodes are in the todo. gpu0–gpu2replace the k8s VMs oncarnaval0–2— k3s servers stopped (onboot 0), then destroyed the same day for the template rebuild; three GPU lab VMs stamped with the roster (history). Kubernetes phase 2 stays open (ops todo → Carnaval).- The
ubuntuaccount's keys come from the roster —role-rack-admins×ssh-keysviaprovisioning/pve/admin-keys.shin the three carnaval scripts, pushed to the running guests (account model).
2026-09-14
- Rack mail fans out to a group address — every consumer now writes to
rack-admins@isc-vs.ch, expanded onsrv-mail, so who gets alerted is one line (ADMINSinprovisioning/mail/deploy-relay.sh) instead of a field per device; Yacine added (group address). Applied and verified on the relay, rumba'srootalias and PVE endpoint, PBS and the iDRAC9 — the PDU and the NAS still name a person (ops todo → Email). - PBS notifications had never been able to send — the
rack-relayendpoint carriedsrv-mail.calypso, NXDOMAIN from that VM. Corrected tosrv-mail.isc3byprovisioning/mail/pbs-notify.shand proven on the relay (PBS).
2026-09-11
- Core Web Vitals of the four main sites — the Umami tag carries
data-performance="true"on the hub, the landing page, this documentation and ISC Learn: LCP, INP, CLS, FCP and TTFB per page, from real visits (service). Heatmaps and replays left off, they would put 190 KB of rrweb on the pages being measured. - Saved Umami reports and a composed dashboard — funnel, goals, attribution and UTM on the
landing page, journey and retention on the hub, retention on Learn;
stats.isc-vs.ch/dashboardshows the four main sites on one screen. Both written by script (provisioning/stats/reports.py,dashboard.py), service.
2026-09-10
- ISC Learn counts with Umami, Google Analytics removed — tag in Additional HTML → head,
the theme's
ganaidga4emptied, caches purged, no downtime (setting, service). - Audience statistics for the ISC web sites — a self-hosted Umami on
srv-stats(VM 121 on rumba,stats.isc-vs.ch), tracker tag deployed on all nine sites, campaign short linksisc.hevs.ch/go/<slug>in the hub (service, site list).srv-freeipa(VM 106,.173) found undocumented on the node while allocating the address.
2026-09-09
- Social-preview cards on the light tone — the picture Teams and Slack show for
docs.isc-vs.chandisc.hevs.chwas dark; both are regenerated light, and the hub's card finally printsisc.hevs.chinstead of the old GitHub Pages address. New filenames (og-card-light.png,og-image-light.png) because Teams caches the preview per URL; recipe inprovisioning/web/og-cards/. - ISC Learn frontpage: ambient background, 44 px rhythm, menu with « ISC Hub » at root — visual pass with pmudry, six deploys (log).
- ISC Learn: theme AJAX blocks requested at HTML parse time — frontpage tabs, blog, « Mes
cours » and the course sidebar panel are fetched by two inline scripts instead of waiting for
first.js; 32 s → 4.5 s under mobile throttling; the theme-file patch for the course panel is retired (log, how it works). The syntax highlighter filter serves its highlight.js files locally instead of from cdnjs (setting). - ISC Learn: 2026-27 students registered — 33 accounts and 371 per-module enrolments from the secretariat's workbook, by the new scripts; per-student enrolments replace the year cohorts (log, procedure).
- Learn toggle sidebar drawn right at first paint, frontpage typewriter script repaired — both CSS/JS-side fixes on hannibal, no downtime (log).
- Docker Hub pull-through registry for the k8s cluster —
srv-registry(CT 120 on rumba,registry.isc3:5000, 200 GB cache not backed up); the three k3s servers pulldocker.ioimages through it, new VMs get the mirror config from cloud-init. Asked for by the Kubernetes course block starting 2026-09-17 (page).
2026-09-08
- Roster groups renamed by kind, operators trimmed to three —
vpn-carnaval/vpn-rack-operators/vpn-rack-mgmtare routes,role-rack-admins/role-pve-auditorare rights;vpn-rack-operatorsnow holds pmudry, yacine.said, adrien.reynard. On NetBird the k8s rangecarnaval-inframoved intonet-carnaval-guests, so everyvpn-carnavalmember reaches the API and Headlamp; the hosts policy folded intooperators-access. Peer login expiration was dropped to 1 h for the switch and put back to 48 h (access model, roster, your access, log). - Keycloak group
calypso-usersdeleted — last remnant of the Calypso model, no consumer since the 2026-08-28 rebuild. - NetBird updated to 0.78.1 (dashboard v2.92.0, agent 0.78.1) over snapshot
pre-0-78-1(log). carnaval8's NVMe hung, GPU lab moved tocarnaval9— the second Swissbit card suspended its pool on 2026-09-03 and frozegpu8for five days; recovered by cold power cycle (poolONLINE, scrub 0 errors), the lab rebuilt asgpu9(VM 1169) andcarnaval8powered off pending the decision on its card (incident, log, todo).- A lab VM that changes node now gets that node's name — repointing an alias hands everyone a changed SSH host key; standing rule on the labs runbook.
- Node-to-node SSH on carnaval explained — the near-empty
/etc/pve/priv/known_hostsis a pre-8.2 vestige, not a broken cluster; PVE 9 uses a per-node file plusHostKeyAlias(access). - marcellus backed up by hand before decommissioning — SSH re-opened by Infomaniak (443 still
closed); every web root,
/etc(with its etckeeper history), logs, homes, the quarantine and InfluxDB archived, plus dumps of all MySQL, PostgreSQL 12 andtisc-dbdatabases, 5.2 GB checksummed on pmudry's machine (C:\temp\marcellus_backup). Enough to rebuild a site elsewhere from its dump anduploads/; core, plugins and themes to be reinstalled fresh, not restored.
2026-09-07
isc-admVM (104) brought into service — hosts the ISC registration manager onhttp://192.168.88.171:8000over the VPN as a systemd unit,just deploypulls and restarts; documented in the rumba guest table. Internal DNS name still to add (todo).- Yacine promoted to NetBird admin — a federated user lands as
userwhatever the roster says; the role is a third act afterrack-adminsandADMINS=, now in the enrolment steps withprovisioning/netbird/set-user-role.sh. - marcellus flagged by Infomaniak and cut off — Shadowserver sinkhole hit traced to a
PyInstaller payload run by
www-data, entry through the Duplicator installer left in the ikarus WordPress since 2021 and used on 2026-09-02; payload quarantined, ikarus disabled, host to be discarded (incident, containment log, plan in the todo). - hannibal hardened the same day — three accounts and their keys removed,
yacine.saidadded as admin, sshd forwarding and root login off, Docker ports closed, 308 updates, automatic security updates back on; afternoon: ingegamez WordPress on auto-updates, Moodle code tree read-only for the web server. Kernel reboot pending (log, state on the server page).
2026-09-06
- SSO admission gate generalised to every service — a login that will not be admitted is now
refused in the realm with a sentence, on every client and not just NetBird; two rules, roster for
machine access and
iscaffiliation for the published tools, withhevsdropped so the rest of the HEI counts as external there (exceptions go throughroster.csv). Both oauth2-proxy gates moved tosession_cookie_minimalafter the new realm role pushed the session cookie past Stirling-PDF's 8 KB header limit (the gate, the user gate). - ISC Learn plugin updates — the six updates Moodle had been reporting, rehearsed on the mirror, 39 s on hannibal, no downtime (log).
- New ISC Learn frontpage in production — ISC-sites visual language on the Moodle landing page, Custom CSS plus a builder page, nothing in the code (log, design).
- ISC Learn content touch-ups — teaching-team page, « Infos » menu, phone pass, course images
for 200.3, 304.2 and 305.2; scripts with
--revertinprovisioning/learn/(log). - ISC Learn mirror refreshed — VM 107 on the day's dump and files; two fixes to the refresh script (log).
2026-09-05
- ISC Learn performance pass — InnoDB buffer pool 128 MB → 2 GB, OPcache sized for the 5.2 tree, event log cut to one year, sessions and application cache moved to Redis; all online, no downtime (log).
- ISC Learn upgraded to Moodle 5.2.2+ — rehearsed on the VM 107 mirror, done on hannibal at 23:15 with 118 s of downtime (log).
- ISC Learn mirror refreshed, rumba copy versioned — VM 107 on the 2026-09-04 dump after a
37 GB delta pull; the copy moved to its own dataset
hdd/hannibal-mirror, ZFS-snapshotted after every pull (log). - Frozen Aug-2024 Moodle removed from marcellus — 81.5 GB of files and the disabled vhost, no
archive (every course is on hannibal);
/srv96 % → 9 %, databases left for a later drop (details). - Moodle course backups purged on hannibal — 800 of 1 194 automated backups removed, retention 3 / 1 by course age; ~92 GB back once the pre-purge snapshot goes (log). Closes the "shrink the Moodle course backups" item.
2026-09-03
gpu8oncarnaval8— node powered on over Redfish (it already held templates9008/9108from the 2026-08-28 restore), VM1168stamped from the CUDA template with the A2 attached, namesgpu8/lab-1168; whole roster aboard,nvidia-smisees the card.carnaval-guests.shnow publishes tocarnaval8–10too (labs runbook).- Yacine's Proxmox access checked —
Administratoron both clusters was already in place; the node Shell asking for alogin:is PVE's root@pam-only rule, now noted on the rumba page.
2026-09-01
- Yacine added as administrator — roster line
staff;rack-admins;appliances-users;mgmt-users;carnaval-users, UID 1105 with a NAS home, root key on rumba, the carnaval cluster andsrv-docker01, PVE user on both clusters, write access onISC-HEI/dc-isc(enrolment, root SSH). The same run restored Adrien's root key on carnaval, lost in the 2026-08-28 rebuild. gpu-01specced — the EPYC 7443 listing of the G242-Z11, 3 × RTX PRO 4500, 8 × 32 GB RDIMM, ≈ €14 000; three dual-slot cards leave a Gen4 x16 for the 100 G NIC, which closes the slot-budget question (architecture §9).- A third R282-Z92 recorded as a cold spare — bought new in August 2026 outside ServerShop24: 2 × EPYC 7302 (Rome), 24 × U.2, no RAM, no drives. No spare CPU needed; its untested backplane is a possible fix for the two defective ones (Epyc, register, todo).
2026-08-30
- VPN session window 24 h → 48 h — NetBird
peer_login_expirationat 172800 s; the deprovisioning tail doubles with it (log). - CRS520 ordered (
sw-100g); the ConnectX-4 Lx count covers every 25 G link, so the NIC line leaves the buy list. UPS candidate set: Eaton 9PX 8000i 3:1 on the 32 A way (architecture §7).
2026-08-29
- The two R740xd are delivered — RG-789539: 2 × Gold 6244, 32 GB, BOSS 2 × 240 GB M.2 boot,
57416 rNDC each, plus 6 PM983, 36 carriers and 3 spare PSUs. Not racked, not yet assigned to
pbs-01/filer-01(register). pve-03accepted and installed —epyc3, the R7515: its universal bays are cabled (12 U.2 slots enumerated), the chassis has four PCIe x16 slots and not two, iDRAC9 is Enterprise-licensed, and the 16 HPE SAS are 512-byte sector. 256 GB fitted from theepyc0DIMMs; PVE installed from a USB key. The 300 GB SAS pair on the order sheet was never delivered, sorpooltakes two of the 1.8 TB (todo, architecture §3).- 16 DIMMs harvested from
epyc0— it was at 1 TB in 32 slots; the 16 Micron modules in theDimm0slots go topve-03(8),epyc1(6) and the shelf (2).epyc0keeps 512 GiB, now at 3200 MHz (log).
2026-08-28
- CIFS sudoers fix written, as a PR (audit F-CODE-1) — branch
fix/cifs-sudoers-pingives the per-usermount -t cifsrule in the Ansiblecommonrole a fixednosuid,nodev,noexec,credentials=~/.smbcredentialsoption string instead of a free*. Not onmain: the role only reaches the CALC machines, which Gregory owns. - Condensation alarm widened — reference raised from 16 to 19 °C (warn ≥ 17, crit ≥ 19): the dew point reached 17.3 °C without any condensation, and 16 fired on ~10 % of the hours (thresholds).
- k3s teaching cluster on carnaval — three GPU VMs on
carnaval0–2, APIk8s.isc3, rebuilt by one script;carnaval-power.shnow re-quorates a partially powered fleet. Operation, what runs. - Calypso retired, every R630 is a carnaval node — the eight Calypso nodes and the three old
carnaval nodes wiped and reinstalled as a fresh ten-voter PVE cluster (
carnavalN= ex-calypsoN, address kept;carnaval5held off, dead DIMM). Local homes rescued torumba:/hdd/backup/calypso-preinstall/first. Templates, CUDA image, SSO and status page redone; k3s is next. Operation, current state, install traps. - Users can no longer change their own email — a review of the day's changes found that a
guestspassword could be turned into any future edu-ID identity: set the guest's address to a colleague's, and their first edu-ID login would be auto-linked into the guest account, groups and all. The address is admin-only now; nobody needed the write. Reasoning on Keycloak. The rest of the review held: guests carryguestsand nothing else, novpn-access, the VM is key-only SSH with two ports open. - Brute-force protection and a password policy on realm
isc— off until now, which mattered little while every login was brokered to edu-ID and matters since theguestsaccounts authenticate with a password on a public endpoint. Temporary lockout only, 10 failures, 15-minute cap; policy is length-only because the passwords are generated, not chosen. Verified live: the counter increments and a legitimate login still passes. Settings and the clear-a-lockout command on Keycloak. - The realm records who changed what — admin events were off, so a change to a group
membership left no trace; they are on now, with representations. Login events were kept 7 days,
now 90. Asserted by
provisioning/keycloak/realm-audit.sh; what is stored and where to read it on Keycloak. The realm's brute-force protection and password policy were still off at that point; see the entry above. - One login page for everything — the accounts without an edu-ID moved out of the user gate's
htpasswd file and into realm
isc, in the new groupguests. That route authenticated before Keycloak, so it carried no password policy, no brute-force protection and no trace in the realm's logs, and it forced a second sign-in page in front of the themed one. Both gates now go straight to Keycloak. Managed withprovisioning/keycloak/local-users.sh; details on Keycloak.
2026-08-27
srv-docker01replaces the static-site container ontest3.isc-vs.ch— VM 105 on rumba, Debian 13 + Docker, root by key for the rack administrators. The proxy forwards the public name to port 8080 of the VM and nothing else, behind the user gate for now. Page: srv-docker01.- Building 19 room surveyed — three three-phase sockets (1 × 32 A, 2 × 16 A), a distribution board in the room, and an existing condensate drain. The 32 A way reopens the UPS input choice that August's 2 × 16 A had closed, and E1/E2/E4 become existing rather than requested. The implantation plan and its printable PDF are at version 1.1.
2026-08-22
isc.hevs.chserves the hub instead of redirecting to it — the address bar keeps the institutional name;/learnis unaffected (log, state).- The hub dropped the
#from its URLs — history-mode routing, so pages areisc.hevs.ch/etudiant-es. Both hosts have the fallback that needs:FallbackResourceon hannibal, and the build copiesindex.htmlto404.htmlfor GitHub Pages, which serves it with a 404 status but the right content. Links of the old/#/…form are rewritten client-side before the router reads the location, so they still land on the right page. The fleet links moved toisc.hevs.chat the same time, making it the canonical of the two names.
2026-08-21
- The root of
isc.hevs.chnow 301s to the hub, which answers on its own namehub.isc-vs.ch;/learnis untouched. Closes the hub-at-the-root plan (log). - Three utility tools for students and staff — Stirling-PDF, IT-Tools and CyberChef on CT 117–119, the first user-facing wave off the service ideas list (service page). All three live on their public names. Only the PDF tool takes a login: the other two compute nothing server-side.
- A second oauth2-proxy instance, the user gate —
iscorhevs, both claim-derived, so it needs no roster. The unit is now a systemd template and the admin gate runs asoauth2-proxy@adminfrom the same files. - The break-glass WireGuard peer drops 40–95 % of its return packets — reported as slow web browsing on an admin laptop; the cause was DNS going into the tunnel over a link losing half its packets, in dead windows of about 14 s (incident). Fixed for the reporter by moving to NetBird; the peer itself is still broken, open in todo → VPN and identity.
2026-08-20
- Scenario 3 is retained, the two R740xd are ordered, and
calypsomastertakesmgmt-01— the scenario page is promoted to Architecture & execution (execution plan now in French), with a new executive summary on top and scenarios 1, 2, 4 archived under considered scenarios. The watcher role moved fromrumbatocalypsomasterfor its iDRAC9 (why); rumba goes to lab/provisioning duty once emptied. Closes the off-site-PBS chassis decision in todo → backups.
2026-08-19
ikarus.snowmon.chfixed — the alias was served over TLS without being on the marcellus lineage, so every visit by that name failed verification. Re-issued with it as a 15th SAN (dry-run first), valid to 2026-11-17; it now 301s to the canonical name. Turned up thatcertonlydoes not reload Apache despiteinstaller = apache(details).- The two worst wiki imports are rewritten — the ISC Learn DRP is now a decision table of failure mode → restore path → duration, folding in what building the standing mirror taught (the delta-pull staging copy, and the mail muzzle that has to be removed deliberately on a cutover). ISC Inventory is 234 → 59 lines and its 3.2 MB of Snipe-IT screenshots are deleted; the UI's own labels carry the procedure.
- Docs clean-up, third pass: the todo — 753 → 654 lines with nothing dropped; the sections grouped by provenance are dissolved into subject ones, two new ones (fabric, DRP) hold what the duplicates asked for, and the reference data moved to the pages that own it (Calypso).
- Docs clean-up, second pass: structural noise — Snipe-IT moved to
Services (on-site),
drp.mdfolded into the ISC Learn plan, both with redirects; journal entries recording sidebar churn, dead wiki template stubs and five closed CALC items are gone. - Audit finding F-SEC-3 closed: no repo file is published any more — a Markdown link to a file
outside the docs tree still emits it as a build asset, and three were live, including
provisioning/uid/uid-map.csvwith 136 rows ofuid,gid,name,email.build/assets/files/is now empty on a clean build; the rule is inCLAUDE.md.
2026-08-18
- Docs clean-up, first pass: the false facts — 26 files. Six pages offered the legacy WireGuard
as a live access path, including the UPS discharge runbook. Also: rumba's model in the
RACKexport, the node and R630 counts, NAS Volume 2 and the home count, and the "RouterOS 6 EOL" case against the CRS326 in three places. - The
sso.isc-vs.chbreak-glass no longer admits the LAN/24— the matcher is NetBird-only (100.65.0.0/16) now, closing the audit finding that the legacy-WireGuard src-NAT let every legacy peer reach the Keycloak console ungated; the stale100.79range was corrected in the same edit. Deployed and verified. What changed for recovery: break-glass. - Both MikroTiks' armed sniffer filters cleared (CCR
vpn_calypso, CRS326ether10–14) — the last piece of the audit's leftovers finding; thewifx.chL2TP client had already gone 2026-08-07. Verified through Oxidized's same-hour pulls. - NetBird admin console gated — the
vpn.isc-vs.chdashboard now sits behind the admin gate; the agents' API, gRPC, relay and the embedded IdP stay open, break-glass mirrors the sso vhost. Verified from outside on all four paths. The gate's sign-in stays onsso.isc-vs.ch, so noproxy_prefixchange was needed. - PBS backups verified after two weeks — every 03:00 vzdump OK (two jobs a night), 284 backup
indexes, dedup factor ~38× (7.05 TB logical in 185 GB), daily GC green.
hddre-checked: 59 %. - Three audit code findings fixed in the repo —
followups.expnow feeds passwords tosudo -S's prompt instead of remote argv (F-CODE-2; the two passwords still want rotating), the netdatakickstart.shin both post-install scripts is pinned and sha256-verified (F-CODE-4), andNB_SETUP_PAT_ENABLEDis off (F-CODE-5). - The hardware-purchase items are settled — the five PERC battery kits are on hand (fitting
them is what remains, R630 hardware); the
S+RJ10cold spare and the pbs-01 SFP28 DACs were dropped by decision. - The three dead marcellus vhosts are disabled (
gptisc,influxdb2,portainer— viaa2dissite, configs kept). Checked before touchinginfluxdb2: no legitimate consumer, only scanner noise (inventory); the InfluxDB service and data stay in place. Live sites verified unaffected. - Hasdrubal settled: destroyed — NXDOMAIN on both zones and absent from the August Infomaniak
inventory. The two references in
change.mdanddrp/isclearn.mdnow carry what mattered (the Infomaniak system snapshot is unreliable) without naming a machine that no longer exists. - A second admin address added to the relay allowlist — applied
on CT 112 and verified by SMTP dialogue (250 for it, 554 for an outside address); every relay
consumer inherits it. The list lives in
RECIPIENTSinprovisioning/mail/deploy-relay.sh. provisioning/web/andprovisioning/status/got their READMEs — the poller's frozen CSV column order and the hand-pushed Telegram token file are now written down (IaC item 7); a Carnaval row was added to the machine table.
2026-08-17
- Acquisition register created from the 13 ServerShop24 invoices (2023–2026) — provenance and component tallies per invoice, and the R630 count reconciled at eighteen: purchases.
2026-08-16
.isc3alias domain added — every.calypsoname also resolves as.isc3(41 twin records on the CCR2004, NetBird nameserver group matching both); the 72 legacy WireGuard peers'client-dnsand the deploy scripts'--searchdomainfixed in the same pass. Details: history.- Retired
admins/teachersNetBird groups removed — empty jwt husks left by the identity collapse, referenced by no policy; the sessions that used to re-create them are long gone. - NetBird updated to 0.77.0 — server, dashboard (v2.91.1) and the routing peer's agent; store migrated clean, access model unchanged, a few seconds of interruption per restart. Details and the new detached update path: history.
- Scheduled jobs wired to Healthchecks — six checks (both vzdump runs, PBS sync + verify,
the hdd scrub via zed, Oxidized liveness), hooks installed by
provisioning/healthchecks/wire-jobs.sh; mechanisms in the checks table. NAS-side tasks remain (todo). - Healthchecks + ntfy deployed — CT 115/116, closing the admin wave (4/4). Healthchecks alerts by email and ntfy (both paths verified); wiring the actual jobs is on the todo. Details: healthchecks, ntfy.
- NetBox seeded — rack, 30 devices with their addresses, 7 prefixes and 13 VMs/CTs imported
from the docs by
provisioning/netbox/populate.py; details on the service page, gaps (cables, serials, Ansible) on the todo. - Oxidized deployed — CT 114
srv-oxidized, both MikroTiks pulled hourly into a local git repo as a dedicated read-only account; admin wave 2/4. Details on the service page; the FS S3600 gap is on the todo. - NetBox deployed — CT 113
srv-netbox, internal-only, Keycloak SSO gated torack-admins; first of the admin wave (NetBox → Oxidized → Healthchecks + ntfy) picked from the service ideas. Details on the service page; populating it is on the todo. - Vaultwarden retired — CT 111 destroyed, vhost, Keycloak client and DNS records removed; it never held a shared secret, the secrets plan being SOPS + age only since 2026-08-13 (rumba history).
- Condensation alarm on srv-status — the poller now publishes the dew point and alerts when it nears the chilled-water inlet (16–17 °C measured; warn ≥ 14 °C, crit ≥ 16 °C), on the humidity tile and over Telegram (thresholds). The past week peaked at 17.3 °C dew point — above the inlet.
- Scenario 3: homes filer decided, backup down to one box — the student homes go to
filer-01, a dedicated standalone filer (twin R740xd of the PBS chassis, closing the §4 A/B question), and the pull-synced second PBS box is dropped for a single hardenedpbs-01in 307; scenario 4's shelf quantities halved to match (§4, §5). - Carnaval lab VMs became cohort machines — accounts and root now come from the roster, keys
from an ISC³-only file, and every VM gets DNS names (
lab-<vmid>+ weekly alias);gpu0–gpu2deployed, one per node with its GPU (stack page). - CUDA image pipeline: build once, publish via the NAS —
--cuda-prepare/--cuda-sealwith a manual-tuning window, dump on the newnas-libraryshare as the cold copy (pipeline, library); templates rebuilt with clang/cmake/nvtop. Trap fixed: the build VM is only reachable through the node, the old builder polled it from the operator's machine and concluded it never booted.
2026-08-15
- rumba rebooted — the UPS USB link did not come back (history, incident).
- NetBird was down ~40 min after the reboot — the no-hairpin hosts entry is now persisted in the cloud-init template (incident).
- A system engineer joins ISC³ end August — "bus factor 1" is restated as a two-person team wherever it carried a decision, in all three scenarios; scenario 3's phase 0 gains the two-person inspection sprint that turns the buy list into a number (scenarios).
- The R7515 order and the NIC shelf folded into all three scenarios — hybrid backplane confirmed
as 12 SAS + 12 universal, two 300 GB SAS shipped so a SAS
rpoolcosts nothing, RAM corrected from CHF 400–550 to ≈ €2 350; 6 ConnectX-4 Lx (25 G ceiling) plus ≥ 6 spare ConnectX-6 mean no 100 G card is bought in any scenario. Scenario 3'spve-03takes all 12 PM983 fornvme_pool(10.5 TiB) withbulkdown to 10 disks, base subtotal ≈ €5 750 (scenarios). - The ConnectX-6 are single-port, which breaks scenario 1 §6.2 —
epyc1's BMC gives MT28908/MT4123 ConnectX-6 VPI, one QSFP56 port, Gen4 x16, fw 20.31.23.54, and the spares are the same model. The "port 2 goes back-to-back to the other node" replication link has no port 2; it now takes a second spare card per EPYC node, so still nothing bought but one more Gen4 x16 slot each (§6.2). Scenario 3 is unaffected — one port per node into the CRS520 is all it asks for. - Scenario 2's chassis is not in stock — no Dell 12 × LFF listed at ServerShop24 on 14 August, nor any MD1200/MD1400 shelf, and the 8 TB SATA line it prices is gone as well; meanwhile the 24-bay all-NVMe SKU its §3 dismissed as uncommon is stocked at €1 639 (option 2).
2026-08-14
- Adrien Reynard has root SSH on rumba — his key added to
/etc/pve/priv/authorized_keys. Who holds root on a PVE node is now the git-tracked list inprovisioning/pve/root-ssh-keys.sh(rumba, access register). - Read-only Proxmox is a roster entitlement now —
pve-auditorsgrantsPVEAuditoron/through the PVE grouppve-auditors-isc, the first rights group managed two-way from the roster (SSO on the node, resource groups). Asserted on both clusters — rumba, and carnaval fromcarnaval0, which covers all three nodes. - Adrien Reynard's access corrected —
appliances-users(no reach to the88subnet, so rumba was unreachable over the VPN) pluspve-auditors, both from the roster. He had been put inrack-admins, which grants no reach at all and would have made him Administrator on rumba at his first login; taken out of it in the console, as that group is add-only in the sync.
2026-08-13
calypsomasterandrumbainventoried over SSH and Redfish — both were mis-modelled in the docs.calypsomasteris a Dell R740 with 8 SFF SAS/SATA bays, not the R740XD §5.2 assumed, so it cannot hold the 10.5 TiB second backup copy (calypso, todo).rumbais a Precision 7920 Rack, has every DIMM slot populated, five free PCIe slots and an uncabled ConnectX-6 nobody had recorded (rumba).- Storage & backup option 2 written — two identical 12-bay LFF chassis for
pbs-02and a production filer, with the R7515-into-production variant and its phase-4 deadline (option 2). Nothing ordered. - The storage plan for the whole estate is settled — dropping the 19 TB
tankturns the dead backplanes into an inventory problem; the allocation across every machine is in §5.2. Nothing fitted yet (todo). epyc1checked over end to end — CPU, storage, PCIe, thermals and firmware all clean (nvme_poolat 4.7/7.8 GB/s, no ECC events across 167 GiB written). The missing 32 GB turns out to be a firmware map-out of socket P1's channelP, and enabling soft PPR did not recover it, leaving a CMOS clear and a CPU swap. An unused ConnectX-6 and an unplugged PSU1 also turned up (checkup, Epyc).- The secrets plan drops Vaultwarden — one scheme for everything, SOPS + age over
secretzone/, with a break-glass age key in the sealed offline set; the deployed instance (CT 111) becomes a removal item (secrets management, todo). Nothing migrated yet. - The hub keeps its host out of its name — the
isc.hevs.chroot redirect will point athub.isc-vs.chserved by GitHub Pages, not bysrv-web01as first recorded: nothing to deploy on the rack, and a later move to another host costs one CNAME instead of edits on hannibal and across the fleet (plan). Still nothing executed.
2026-08-12
epyc1documented and its storage built — a Proxmox VE standalone node on192.168.88.46, BMC pinned to192.168.88.12(epyc1-bmc), with its three 3.2 TB NVMe asnvme_pool: RAIDZ1, 5.68 TiB usable, fromprovisioning/pve/epyc1-nvme-pool.sh. New page: Epyc.epyc0is in repair. Credentials insecretzone/epyc.md; root password rotation and the missing mail path are in the todo.- The racked Tango node answers on
192.168.88.43, not the documented.20— its MAC has no reservation (Tango, decision in the todo). - Both EPYC U.2 backplanes are defective, different bays on each, so the PM983 drives stay out
and the 19 TB
tanktarget is blocked. Diagnosis, the power-only LED test that localises it, and the bay/CPLD reference are on the Epyc page; the four options, cheapest first, are in the todo. Also recorded there:epyc1's dead memory channel and the deliberate empty DIMM slots that look like faults but are not.
2026-08-11
- SInf also opened TCP 587 to Infomaniak — verified from rumba; 587 to any other host and 25 to anything still time out, so it adds a port, not a second provider. The relay stays on 465/implicit TLS (Email, uplink restrictions).
- Wording pass over the whole docs tree — flattened PR-voice emphasis in 75 files (bold used
to insist, intensifiers, verdict sentences) without touching facts, headings, numbers or links.
The rule now sits in
CLAUDE.md. - Factual contradictions found by that pass, corrected — node counts (the rack has 8 Calypso
and 3 Carnaval, neither page had it right), the NAS exports being
asyncnotsync, the frozen Moodle's size, NetBird's session window, the CRS326 having been upgraded after all, the WireGuard peer count (74, from the 2026-08-10 config export), two:::closers that swallowed the paragraph after them, and assorted malformed markup. What still needs a decision is in the todo. - UPS replugs failed, NUT still blind — both connectors reseated, neither produced a kernel event while port 13 reports healthy, leaving the cable, the UPS's own port, or a dead port 13 (incident). The unit does have a DB9, so repair-or-move-to-serial is now a real choice rather than a hypothetical.
- NetBird deprovisioning documented properly — a NetBird-side block cuts immediately, disabling the upstream account alone leaves a 24 h tail, and the way to revoke now is to end the session by hand in Keycloak (netbird).
- Decision: the root of
isc.hevs.chwill point at the hub (hub.isc-vs.ch),/learnuntouched — recorded in the learn migration plan, work items on the todo; nothing executed yet.
2026-08-10
- tMatch is gone from marcellus — everything behind it deleted on request without a backup,
which also closed 16 root-equivalent entry points and the FreeIPA-on-
0.0.0.0audit finding. The source is safe on GitHub; the database is not (details). - Marcellus's shared certificate too — re-issued without the NXDOMAIN
apt.chezmoicamarche.ch, 18 names valid to 2026-11-08, all verified from outside; its renewal was due around 2026-09-04 and would have failed all 19 vhosts (marcellus). Both shared lineages are now clean. - hannibal's shared certificate is fixed before it could take seven vhosts down — re-issued
without
wiki.isc-vs.ch, the missing Apache reload hook added, standing rule written (log, rule). - The rack can send email, for the first time — SInf opened 465 to Infomaniak and everything
goes through CT 112
srv-mail, the Postfix relay that holds the only copy of the credential. Six consumers wired the same day,zfs-zedandsmartdamong them, which had been mailing into a void (Email, rumba log). secretzone/moved to the repo root (withsecretzone/calc/for CALC), out of both Docusaurus trees — the published site can no longer expose it even by misconfiguration; all referencing docs and provisioning scripts updated (secrets management).
2026-08-09
- The UPS's USB link wedged for 7 h, blinding NUT from 05:03; a
usbresetbrought both interfaces back with the UPS itself never at risk. NUT'sinsufficient permissions on everythingmeans a wedged link, not a permissions problem (recovery, incident). - The CRS326 is off RouterOS 6 — 6.49.19 → 6.49.20 → 7.21.5 → 7.23.3 in one afternoon,
RouterBOOT matched, so both MikroTiks now run the same current firmware; config conversion intact
(details). Two standing rules
came out of it: a MikroTik upgrade cannot be predicted from npk size versus free space, and v7 is
offered to a v6 device only on the
testingchannel (both). Still open: L3 offload, now within reach (todo). - The CRS326 can reach the world again, after two independent breakages: its DNS pointed at the
retired SInf gateway
172.30.7.1, and its owninputchain ended indropwith noestablished,relatedaccept, so replies to traffic it originated were dropped by itself — theforwardchain does have that rule, which is why everything behind the switch was online while the switch was not. NTP syncs andcheck-for-updatesworks; the rule is on the network page. - The routers have a remote rescue path, and a smaller management surface — RoMON enabled on
both with a shared secret and forbidden on the WAN port, so WinBox reaches the CRS326 through
the CCR when its IP config is broken (exercised: discovery finds it at one hop).
telnet,ftpand plainwwwdisabled on both — part of F-NET-4;api(8728) stays,routeros-api.pyneeds it (todo, audit notes insecretzone/security-audit-2026-08.md). - CCR2004 upgraded to RouterOS 7.23.3 (from 7.20.4) with its RouterBOOT to match — two reboots, under a minute each, config and services verified end to end (details).
- The CCR's pre-WireGuard remote access is deleted — L2TP server (
use-ipsec+ shared secret), thevpnPPP account and thewifx.chL2TP client to the former ISP. All four were already disabled and never used; they only survived as three cleartext passwords in the config (network). - MikroTik backups are now complete — both devices re-exported with sensitive values
(Wireguard private keys, the
wifx.chuplink password, the L2TP/IPsec secret) and their binary backups copied off for the first time, so user accounts and certificates are covered too. One script does all four steps:provisioning/network/mikrotik-backup.sh(process).
2026-08-07
- ISC Learn migration plan written — target: Infomaniak Serveur Cloud managé (proposal); prod facts refreshed read-only (log).
- hannibal's fail2ban now bans — the sshd jail watched port 2002 while sshd listens on 20002;
fixed, and
ubuntugiven a console break-glass password (log). - Marcellus is inventoried — read-only survey of the legacy VPS: 19 vhosts (5 answering 503),
6 containers, ChirpStack, four PHP-FPM versions and a frozen Aug-2024 Moodle using ~81 GB
of a 96 %-full
/srv(Marcellus). Turned up a dated [HIGH]: a name on its shared 19-vhost certificate no longer resolves, which would fail the renewal for all of them around 2026-09-04. Decommissioning otherwise blocks on naming the owners of the non-ISC vhosts (todo). - The
pagodepair is renamedtango— the name fits a two-machine pair, and the old name followed no convention. The planned CALC@HEI login node that already held the name Tango becomes mambo (secretzone/calc/mambo.md). Also placed at U36 front on the schematic, where it actually belongs (Tango). - What a Proxmox cluster brings is now written down — the docs only carried the costs (quorum, join/leave constraints, HA vs. thermal shutdown); new section splitting what a plain cluster gives from what needs shared storage (§4.1), plus cluster/corosync/quorum/ QDevice in the glossary — renamed from "Acronym glossary" to hold terms too. Stale PBS retention (7/4/3 → 14/8/6) and datastore size fixed on the backups and rumba pages.
carnaval1's NVMe controller hung and tookfast-vmdown for 17 h — recovered by a cold power cycle from the iDRAC; the pool scrubbed clean and the card is back in service, on watch (incident, recovery sequence).- The legacy MikroTik WireGuard fleet is off — 12 peers removed and every other one disabled,
leaving
wg13as the only enabled peer. Attribution table, what was removed and the break-glass consequence:secretzone/wireguard.md. - Security audit of the whole architecture — read-only; findings and the ordered remediation
list in
secretzone/security-audit-2026-08.md, open items marked by severity on the todo. Also stopped two files linked from the calc pages being published as site assets.
2026-08-06
- The disposable GitLab test pair is destroyed — CT 105
srv-gitlab(with its secretzone copy and full-rights API token) and VM 106srv-runner01are gone, along with their five DNS entries and the teardown checklist page itself; what the test proved and the deltas for a real instance: history entry. - Two group deletions silently took five VPN policies with them — ~1 h of access-model outage,
restored in one
replay-config.pyrun from the committed snapshot (its first real recovery); the cascade is now a standing rule on the service page (incident). - The identity model collapsed to
students/staff— identity now grants admission only; the oldteachers/adminsreach became the entitlementsappliances-users(appliance subnet- carnaval PVE hosts) and
mgmt-users(iDRACs), so every grant is per-resource. NetBird policies repointed byprovisioning/netbird/migrate-role-groups.py(access model, role groups). Firststaffmember: Marta Rende.
- carnaval PVE hosts) and
- Enrolling is now a documented two-step — roster line + sync, with the un-enroll path and its deprovisioning tail: enrolling people.
- The ISC wiki is retired — DokuWiki archived off hannibal (
rumba:/hdd/backup/wiki-archive/, where it is) andwiki.isc-vs.chturned into a 301 vhost onsrv-web01that maps the olddoku.php?id=…links onto their new pages (the vhost). - The documentation has its own name:
docs.isc-vs.ch— GitHub Pages custom domain (static/CNAME), so the site is served at the root instead of under/dc-isc/. GitHub redirectsisc-hei.github.io/dc-isc/<path>there by itself, so old links keep working. - A non-enrolled edu-ID login to the VPN now gets told why — Keycloak denies NetBird logins
lacking the roster-derived
vpn-accessrole with a themed EN/FR "not enrolled — contact the staff" page, instead of NetBird's bare error after a successful login (the VPN gate,provisioning/keycloak/vpn-access-gate.sh). The first binding broke every edu-ID login for ~20 min — the trap and its fix are a standing rule on the Keycloak page. - Pre-created roster accounts now link silently on the first edu-ID login — the
eduid-autolinkfirst-broker flow replaces a verification Keycloak could not deliver, and edu-ID becomes authoritative for the stored profile. Validated end to end with two students, VPN gate included (details).
2026-08-05
- DHCP is off on the node subnet — the CRS326 pool overlapped the carnaval lab-VM range while
the VMs are assigned statically, and nothing had used the server since the MAAS client that held
its only lease disappeared. Pool moved to
192.168.91.200-.240so a careless re-enable cannot land in.128–.191, both servers disabled (todo → network follow-ups). - Lab VMs get their address and their name by themselves —
carnaval-lab-vm.shallocates the lowest free octet in192.168.91.128–.191from the guests the cluster actually runs, registerslab-<vmid>.calypsoon the CCR2004, and--destroydrops both. The range is the NetBird resource, so a VM can no longer land where no VPN client reaches it (how to run it). - Numeric UIDs are issued from a register, and Ansible reads it —
provisioning/uid/uid-map.csvseeded from the ownership NFS actually compares, withuid-alloc.pyto allocate and verify andcreate-home.shto make a home with the number it holds. Applying it will renumber 13 rack accounts,uid-alloc.py --check-ansiblelists them (UIDs). elia.pacionirenumbered 1027 → 1100 — 1027 isIscAdminin the NAS's own/etc/passwd, so DSM rendered that student home as owned by the admin account. Moved on the NAS and the three nodes together, so nothing drifts apart.- Seven departed accounts removed from the Calypso nodes, their SSH keys and the register, and
the last two
/hometrees dropped from the carnaval pre-install rescue on rumba. Their numbers are free again; nothing else in the 203 rescued homes was touched. - The roster now drives VPN access per resource, not per role.
roster.csvbecame one line per person (email,groups[,nasname]), 18 third-years and 3 teachers are in, and the flat192.168.91.0/24grant was split socalypso-usersandcarnaval-usersmean different things (role groups, access model). - UID bands widened and given a hole for DSM — the old plan ran out after 500 students and collided with the Synology's own accounts (UID allocation).
- The SSO login page now names the service you are signing in to and hides the local
username/password form behind a disclosure — edu-ID is the only thing the card offers at first
glance, which matches reality: realm
ischolds no password at all. Theme inprovisioning/keycloak/theme/(Keycloak page). - New index: service access paths — the URL or address of every service and what it asks you for, in one table set, linking to the page that owns each one.
- Proxmox now takes an edu-ID login through Keycloak, on rumba
and on the carnaval cluster — realm
isc(clientproxmox),root@pamuntouched and still preselected. Only pre-created users are admitted (--autocreate 0), and the sole Administrator grant on each cluster is the PVE grouprack-admins-isc, filled from the Keycloakgroupsclaim at every login — so revoking there revokes here. PVE appends the realm name to claim groups, one of the traps recorded on the node page. - The recurring netdata UPS mails were a false alarm, now silenced at the source —
upsd_ups_last_collected_secsflapped 3–4×/hour (40 pairs in two days) on stock's 5 s threshold, which the driver's stale-data spells cross routinely; it now warns at an absolute 60 s (why, and why the collector must not be reconfigured). The mail path was Netdata Cloud — rumba is claimed, and its own postfix has delivered nothing in days. - Carnaval nodes added to the rack status page, with a netdata deep-link — the poller was still
calling them
calypso8–10; any machine card flaggednetdata: truenow links straight to its dashboard (how). - No lab VM mounts the whole
homestree any more — the pre-2026-08-05 VMs are gone (the cluster holds only the six templates), so that item closes; the same exposure on the Calypso nodes, wheredockermakes students root-equivalent, stays open until the rebuild. --sudogrants root, which it did not before —useraddleaves the account with no password at all, so group membership alone leftsudoprompting for a credential that cannot exist.carnaval-lab-vm.shwrites aNOPASSWDdrop-in per granted name, the way cloud-init already does forubuntu(how to run it).- A GPU guest reads as 100 % memory used, and that is expected — VFIO pins every page and
passthrough forces
balloon: 0, so PVE reports the qemu process RSS rather than guest usage: 62 GB against 1.3 GB actually in use (why). Same pass corrected two stale claims —ubuntu-2404-cudaexists on all three nodes, and the reference guest VM 110 is gone, so all three GPUs are free. - Lab-VM UIDs now come from the NAS, not a login node —
carnaval-lab-vm.shstats the home directory, because the ownership NFS enforces is the only source that cannot be stale: a Calypso node'spasswddisagreed for 4 of 67 accounts, and a wrong number fails silently. It also refuses a home owned byuid 0. The permanent fix — issuing UIDs once instead of rediscovering them — is designed in UID allocation; the four drifted accounts and the unusedgroupe1–groupe4dirs are open items. --sudotakes names —--sudo=pmudrygives root to one person on a lab VM instead of the whole class, so a teacher can hold a privileged account of their own with their real UID and NAS home (how to run it).- Two home-ownership repairs on the FS2500 —
olivier.amacker's home was ownedroot:root, restored to his10011:10011; the four unusedgroupe1–groupe4project directories and their groups oncalypso0/1/2are gone (open items). - A NetBird restart destroyed its own store —
encryptionKeywas empty, so the server had been minting a throwaway key at every start; ~90 min VPN outage, carried by the legacy WireGuard, and the store rebuilt fromconfig-snapshot.jsonwith the newreplay-config.py(incident, what the account looks like now). - The VPN logs in through Keycloak, and access follows the roster —
netbird upgoes to ISC SSO then edu-ID, and admission is membership of a role group instead of a manual approval nobody could see (how it works, what it took). Remaining: open items. - edu-ID logout no longer ends on a French error page — the provider's logout URL is now empty, so logout ends at Keycloak and returns to the client (why).
- NetBird lazy connections off — they were stranding a connected peer with correct routes; relay-only with three peers gains nothing from them (detail).
- History sub-pages made uniform — one name for all of them:
<page>-history.md, titled "history & operations", newest entry first.nas-operations.mdwas renamed and the rumba page's sections flipped to newest-first; the convention is now in thedocument-workskill. - Ansible lives in this repository —
provisioning/ansible/, playbooks and their configuration together, laid out the way the setup script used to assemble them so no playbook changed. The two GitHub repos are archived. Why: the roster and SSH keys sat in one repo while everything consuming them sat elsewhere (how it works now). - Lab VMs no longer expose the whole cohort's homes — one NFSv4.1 subdirectory mount per student
instead of the
homestree, andsudois now opt-in (--sudo), since root on any box in the node subnet can impersonate any UID oversec=sysNFS. Verified: a student reads their own home, another student's path does not exist (how to run it). - The MikroTiks are driven over their API, not SSH —
provisioning/network/routeros-api.py. RouterOS echoes log lines into interactive sessions (stockaction=echo topics=critical), which races every scripted pattern-match against the ~10 s login timeout; the API has no terminal and answers in under a second. Used it to disable the Dude server, the source of the constantlogin failure … via winboxentries (it polls with credentials that no longer match). - Student-home NFS mounts work again — the 2026-08-02 hardening
had made every new NFSv4.1 mount of
calypso_homes/homesfail while the live ones kept working, so a node reboot would have lost its homes. Fixed with one non-inheritingr-xACL entry forgueston the share root, keepingroot_squash: what it was and why that fix. - NAS writes went from 65 to 102–109 MB/s — both exports switched to
async, which was the bottleneck; the client-sidesynceveryone assumed was the problem costs nothing on top (measurements and the durability trade-off). Writes now match reads at 1 GbE wire speed. Also corrected on the way: the claim that the FS2500 refuses NFSv4 was wrong.
2026-08-04
- Carnaval has an operator runbook — Running labs on carnaval:
handing a VM to a student, attaching a GPU, cohort loops, end-of-semester cleanup, and rebuilding
the cluster from the repo. Also on 2026-08-04, the duplicate
dhcp_testserver on the CRS326 was disabled, so a VM can no longer be handed a node's address. - Carnaval can hand out lab VMs, GPU included — base and CUDA templates on the three nodes plus the cohort pools; clone to SSH in ~30 s, and a T4 passed through ran a CUDA kernel clean. Conventions, the VMID/address plan and the passthrough recipe: cluster page. Remaining: todo.
- The playground cluster exists —
calypso8–10wiped and rebuilt as the quorate 3-node Proxmox cluster carnaval, which starts phase 3; the old/homewas rescued to rumba first. Two traps written up where they belong: the BIOS cannot boot the NVMe cards, and iDRAC8 needs its own install path. - PSU criticals cleared on
calypso0–calypso3, and the R630 fleet swept for PERC batteries — an iDRACGracefulRestartre-inventories the chassis and drops a deliberately absent PSU (Critical→OK); clearing the SEL alone does not, and it never clears an absent BBU. Three nodes (calypso1,calypso2,carnaval0) are missing theirs and stayCriticaluntil the kits arrive — table, order and traps in ops todo → R630 hardware. - Admin panels are behind a login instead of an IP range, and the split-horizon DNS layer is
gone —
oauth2-proxyonsrv-web01guards the admin paths against the Keycloak grouprack-admins, so they work from anywhere with no VPN and no client-side DNS trickery, and the layer that forced an internal answer for a public name is deleted. A LAN/VPN break-glass branch stays, since the gate authenticates against Keycloak itself (how it works). - Password vault live, and it is Keycloak's first consumer — Vaultwarden 1.37.1 as CT 111
(
srv-vaultwarden), published atvault.isc-vs.ch, login through the realmisc(SSO_ONLY, invite-only). Keycloak login does not replace the per-user master password. (Retired 2026-08-16 without entering service.) NeededtrustEmail=trueon theeduidprovider, otherwise brokered users arrive unverified and Vaultwarden refuses the signup. - Rumba's own configuration is now code —
provisioning/pve/post-install.shcovers everything the installer leaves out: deb822 repos, thehddRAIDZ2 pool, the five storages, both nightly backup jobs, netdata. Idempotent, so it doubles as a check that the node still matches its page. Closes the first of the IaC audit gaps; items 4–7 there remain. - Keycloak realm and NetBird access model are snapshotted into git —
export-realm.sh→realm-isc.jsonandexport-config.sh→config-snapshot.json, secrets stripped, to be re-run after any console/dashboard change and committed. Both are diff tools, not restore paths: neither API can import what it exports — Keycloak, NetBird. - edu-ID federated login works and feeds three institution groups — the resource registry
approved
hes-so_isc3_vs_oidc_sso, a real login brokers a user, andhes-so/hevs/iscare filled from the claims (how). Marking an ISC member took an RR amendment; the measured claim values and the four traps are insecretzone/eduid-oidc.md. - SSO login page now wears the ISC brand — Keycloakify theme (petal colors, logo,
dark mode) built in
provisioning/keycloak/theme/, deployed withdeploy-theme.sh,loginTheme=iscon the realm. Build/deploy notes on the Keycloak page; provider jars now survive in-place upgrades. - Rumba's guests now back up off-host — PBS 4.2 as a VMM VM on the FS2500, nightly job for all
guests, chain verified with a live restore-capable backup; Volume 2 became PBS-only and the local
job was trimmed to
keep-daily=3sohddstops overflowing. The rationale and the VMM/token traps are on the service page.
2026-08-03
- Netdata's stock UPS battery-charge alarm silenced — it keys on
nutdrv_qx's unreliable charge estimate, and the UPS firmware intermittently zero-fillsbattery.voltagein its Q* replies, so the alarm fired with the UPS on line. Override now installed bydeploy-ups.sh; details. - Rumba moved onto the rack UPS (both PSUs, ~5 s per swap, no reboot). UPS load 14 % ≈ 265 W — rumba is the only thing on it, and the UPS input is itself a PDU outlet. Measurements in the rumba history. Still open: NUT is monitor-only, so a long outage hard-kills rumba — todo.
- UPS status on the status page — NUT polled every 30 s,
UPS tile + load chart + history columns, Telegram alerts keyed on
ups.statusonly (never on the estimated runtime). - UPS real battery telemetry decoded — the second USB interface (HID Power Device) reports a
measured autonomy of 55 min at rumba's load; NUT's own charge/runtime estimates are unreliable
on this unit. Read with
ups-read-hid; details and traps in the HID section. - NAS given one dedicated 10 GbE port per subnet — second cable,
eth3static192.168.91.250/24; rumba ↔ NAS measures 1.1 GB/s. Bonding rejected on purpose. NAS history; rebuilt nodes must follow how a node must mount it. - UPS monitored over USB (NUT) — the BlueWalker has no network interface, so USB from rumba is
the only management path. Service page; runnables in
provisioning/ups/. - Rumba's update backlog cleared — 0 pending,
pve-manager9.2.6 / kernel7.0.14-8-pve. - Docs reorganized into present / past / future tiers — machine/service pages hold current state only; dated logs moved to history sub-pages; incidents in the incident log; open items consolidated in the todo.
- NetBird masquerade turned off — appliances now see each user's own overlay address
(
100.79.x.x), so logs and per-IP protections act per user; the July 2026 lockout pattern is gone on this path. Service page. - Keycloak identity broker deployed — CT 110,
https://sso.isc-vs.ch, realmisc, admin console LAN/VPN-only. Service page; runnables inprovisioning/keycloak/. - SWITCH edu-ID OIDC client registered —
hes-so_isc3_vs_oidc_sso, awaiting RRA approval. One client for the broker; GitLab, Grafana and Moodle become Keycloak clients with no further AAI paperwork. Record and constraints in the secretzone (eduid-oidc.md). - NetBird: edu-ID needs no licence — external OIDC login and JWT group sync are Community features; only SCIM directory sync is Enterprise. Do not buy it. Details.
- Proxmox web UI given the ISC typography — self-hosted Manrope/Inter, idempotent deploy with
--revert. Runcheck-coverage.shafter everypve-manager/extjsupgrade. Rumba history. - NetBird group profiles validated client-side — admin and student profiles match the access model exactly, split tunnel confirmed. NetBird history.
- NetBird deployed in production —
vpn.isc-vs.ch, VM 109, everything on the existing 443; relay-only is the permanent mode (the uplink drops inbound UDP 3478). The MikroTik WireGuard stays as the emergency/admin path. Service page.
2026-08-02
- Network performance audit run and fixes applied — node↔node is wire-speed 1 GbE; the
91↔88path was software-routed at 38.7 MB/s, fixed with on-link routes (61.5 MB/s NFS writes), afasttrack-connectionrule and thelayer-3-and-4bond hash policy. No new hardware needed. Network history; resulting config in the storage fast path. - Calypsomaster audited (read-only): nothing load-bearing runs on it. Findings and the conversion-to-second-PVE-node plan in the todo.
- Ansible repos clarified:
ansible-playbooks(+ private-conf) are the current fleet management; the copies on calypsomaster are stale duplicates of the archived repo. - FS2500 given its own NAS page with the first disk-level readout — Volume 2's Toshibas are 7.4 years old, older than the NAS itself.
- FS2500 NFS exports hardened — three real holes closed (a
/24mask that admitted every VPN user as root,no_root_squashon the node rule, delete rights on the homes ACL). Details. - FS2500 data scrubbing run for the first time (Volume 1 had never been scrubbed in
1.8 years) — 0 errors,
mismatch_cnt=0; quarterly schedule enabled on both pools. Details. - FS2500 Volume 2 rebuilt from 7-disk RAID 5 to 6-disk RAID 6 + a hot spare covering both pools — survives two failures, and Volume 1 gets automatic rebuilds for the first time. Capacity 11 → 6.7 TB, at no cost (the volume was empty). Details.
- FS2500
rumba_backupstub removed — 6.7 TB of free RAID 6 available for the DS923↔rack replication goal (still in the todo).
Earlier in 2026
- Patch panel for room 23N307 searched and bought (server-room list). Placement constraint to respect at install time: outside the rack, not inside — rack space will be premium.