Incident operating rhythm
Coordinate evidence, owners, checkpoints, and communication without turning the tool into another status dashboard.
Establish the clock
At the start of an incident, record the impact window, the next evidence checkpoint, the next decision checkpoint, and any external communication commitment. A checkpoint should have an owner and a concrete expected output. "Update in thirty minutes" is weaker than "Storage owner returns affected-versus-healthy latency evidence by 14:30."
Keep four records distinct
| Record | Purpose |
|---|---|
| Evidence ledger | What was observed, where, when, and by whom |
| Decision log | What the team decided and the evidence used |
| Action record | What changed, who authorized it, and expected outcome |
| Communication record | What was shared, with whom, and the next commitment |
- Record
- Evidence ledger
- Purpose
- What was observed, where, when, and by whom
- Record
- Decision log
- Purpose
- What the team decided and the evidence used
- Record
- Action record
- Purpose
- What changed, who authorized it, and expected outcome
- Record
- Communication record
- Purpose
- What was shared, with whom, and the next commitment
Run the narrowing loop
At each checkpoint, ask what changed in the evidence state. Remove boundaries that are no longer supported. Promote boundaries only when evidence supports them. Assign the next unanswered question to a named owner. Avoid reopening settled questions unless new evidence contradicts the previous conclusion.
Close responsibly
Closure requires impact recovery, technical validation, owner agreement, and residual-risk recording. If the immediate service is restored but the underlying boundary remains uncertain, close the active incident only with a separate problem investigation, owner, and due date.