Chapter 12 · operating safely

Technology domain walkthroughs

Apply the same operating logic across HPC, AI platforms, cloud, networking, storage, and data-center operations.

2 min read·Updated 2026-07-24·2 role paths
01

HPC and AI infrastructure

GPU, fabric, scheduler, thermal, memory, storage, and workload symptoms cross ownership boundaries. Define the affected job population, compare healthy nodes or queues, and separate workload behavior from shared platform behavior before assigning a hardware or scheduler boundary.

02

AI platform

A model-serving symptom may arise from model quality, capacity, dependency latency, policy, data, orchestration, or request shaping. Preserve request class and healthy-path comparisons. Do not label a quality issue as infrastructure without evidence, or an infrastructure issue as model behavior because the output is visible at the application layer.

03

Cloud

Cloud symptoms can sit in identity, quota, network, service control plane, workload configuration, regional dependency, or provider service. Use region, account, request class, and dependency comparisons to avoid treating the first visible service as the fault boundary.

04

Networking

Define source, destination, path, protocol, time window, and expected behavior. Separate reachability, loss, latency, policy, name resolution, and application response. A packet path should be paired with control-plane and endpoint evidence when relevant.

05

Storage

Separate capacity, latency, throughput, pathing, metadata, permissions, and application access patterns. Compare the affected mount, volume, or object path with a healthy peer under similar load.

06

Data center and physical systems

Power, cooling, rack position, cabling, environmental conditions, firmware, and maintenance history matter. Physical work requires explicit safety and authorization. Preserve photographs, labels, timestamps, and the exact asset identity.