Technology domain walkthroughs
Apply the same operating logic across HPC, AI platforms, cloud, networking, storage, and data-center operations.
HPC and AI infrastructure
GPU, fabric, scheduler, thermal, memory, storage, and workload symptoms cross ownership boundaries. Define the affected job population, compare healthy nodes or queues, and separate workload behavior from shared platform behavior before assigning a hardware or scheduler boundary.
AI platform
A model-serving symptom may arise from model quality, capacity, dependency latency, policy, data, orchestration, or request shaping. Preserve request class and healthy-path comparisons. Do not label a quality issue as infrastructure without evidence, or an infrastructure issue as model behavior because the output is visible at the application layer.
Cloud
Cloud symptoms can sit in identity, quota, network, service control plane, workload configuration, regional dependency, or provider service. Use region, account, request class, and dependency comparisons to avoid treating the first visible service as the fault boundary.
Networking
Define source, destination, path, protocol, time window, and expected behavior. Separate reachability, loss, latency, policy, name resolution, and application response. A packet path should be paired with control-plane and endpoint evidence when relevant.
Storage
Separate capacity, latency, throughput, pathing, metadata, permissions, and application access patterns. Compare the affected mount, volume, or object path with a healthy peer under similar load.
Data center and physical systems
Power, cooling, rack position, cabling, environmental conditions, firmware, and maintenance history matter. Physical work requires explicit safety and authorization. Preserve photographs, labels, timestamps, and the exact asset identity.