Reliability and operability¶
Evaluation question¶
Can the team detect, recover from, replay, and safely operate failures?
Reliability is not the absence of failure. It is predictable behaviour when failure occurs and a recovery process the owning team can execute.
What to examine¶
- Idempotency and duplicate protection
- Bounded retry and backoff
- Replay and recovery boundaries
- Recovery objectives
- Failure ownership and escalation
- Runbook clarity
- Operational effort relative to team capacity
Evidence in the first Decision Case¶
The retail sales case must include one recoverable processing failure. The engineer must detect it, determine its effect on the reporting outcome, recover or replay safely, and verify that no duplicate or missing result was introduced.