Operations
How ISC³ is run. This section is written for the staff operating the datacenter; users of the service should head to Using ISC³ instead.
Where things go on this site: this section describes what humans do — processes, monitoring, backup policy; what exists and how it is built lives in Infrastructure. Acronyms are expanded in the glossary.
Operations of the research nodes (Chacha, Disco) live in the CALC@HEI section — user management, admin scripts and the research todo list moved there.
Access
- Service access paths — the URL or address of every service, and what it asks you for
Processes
- Change Management
- Incident Management
- MikroTik configuration backup & restore
- Unattended Proxmox install over iDRAC
- GitLab test environment teardown
- Disaster Recovery Plans
Monitoring & backups
- Monitoring & alerting — where to look, in escalation order (rack status page & Telegram bot, Netdata, PDU), what alerts exist, and what to do when the thermal alert fires
- Backups — what is backed up where, retention, and the known gaps
Tooling
- Ansible — configuration management for the whole fleet (shared with CALC@HEI for now)
Inventory
- Inventory
- ISC Inventory — the Snipe-IT instance
Secrets
- Secrets management — where credentials live and who can read them