Skip to content

Reliability and operability

Evaluation question

Can the team detect, recover from, replay, and safely operate failures?

Reliability is not the absence of failure. It is predictable behaviour when failure occurs and a recovery process the owning team can execute.

What to examine

  • Idempotency and duplicate protection
  • Bounded retry and backoff
  • Replay and recovery boundaries
  • Recovery objectives
  • Failure ownership and escalation
  • Runbook clarity
  • Operational effort relative to team capacity

Evidence in the first Decision Case

The retail sales case must include one recoverable processing failure. The engineer must detect it, determine its effect on the reporting outcome, recover or replay safely, and verify that no duplicate or missing result was introduced.