Instiq

6System operation & facilities

Practice questions →Glossary →
  • 6.1System monitoring and job management

    Covers monitoring items (CPU, memory, disk, response time) and the design of thresholds and alerts that catch anomalies early, job scheduling that protects the SLA, the dependencies and abnormal-end handling (abort, rerun, skip) of a job net that ties multiple jobs together, and two-stage warning/critical thresholds that suppress false alarms.

  • 6.2Backup operation, automation, and operating procedures

    Covers operational automation that hands routine work to machines, the runbook / operating procedures that document steps, human-error prevention through double-checks and call-and-response, and safely applying configuration changes (change-management linkage and rollback) to production. The key is separating routine work to automate from exceptions requiring human judgment.

  • 6.3Facility management

    Covers the UPS (uninterruptible power supply) and on-site generator that bridge a momentary outage to standby power, the air conditioning that cools equipment and the power distribution that delivers electricity, seismic/base-isolation protection against earthquakes, access control preventing unauthorized entry, and the data-center Tier levels that stage a facility's availability. The key is choosing the degree of redundancy by trading off the availability target against cost.

  • 6.4Operations design and cost, staffing, and budget management

    Covers operations design that builds the operating structure and procedures, the design of shifts and staffing that runs a 24-hour operation seamlessly, cost optimization that balances investment and running cost, KPIs that measure operational quality numerically, and supplier / resource / budget management that governs external suppliers, resources, and budget. The key is building a structure that is neither excessive nor deficient within the SLA-versus-cost trade-off.