Skip to main content

Operations

How ISC³ is run. This section is written for the staff operating the datacenter; users of the service should head to Using ISC³ instead.

Where things go on this site: this section describes what humans do — processes, monitoring, backup policy; what exists and how it is built lives in Infrastructure. Acronyms are expanded in the glossary.

Operations of the research nodes (Chacha, Disco) live in the CALC@HEI section — user management, admin scripts and the research todo list moved there.

Access

Processes

Monitoring & backups

  • Monitoring & alerting — where to look, in escalation order (rack status page & Telegram bot, Netdata, PDU), what alerts exist, and what to do when the thermal alert fires
  • Backups — what is backed up where, retention, and the known gaps

Tooling

  • Ansible — configuration management for the whole fleet (shared with CALC@HEI for now)

Inventory

Secrets

Open points