Launch is a checklist, not a date. The distance between a system that works in staging and one that survives production traffic is a series of gates — each with sign-off criteria and named approvers.

Our readiness review covers threat modelling with findings fixed, permission and negative-path testing complete, load validation at multiples of expected traffic, release rehearsal with rollback proven, monitoring tuned to the system, and incident rehearsal with the actual on-call engineers.

Sign-off, not optimism

Each gate names its approvers on both sides. Skipped gates are recorded as accepted risks with owners — never as silent omissions. This sounds bureaucratic until the first incident, when it becomes the reason the response takes minutes instead of hours.

Monitoring deserves special emphasis. Balance watchers, anomaly detection and finality alarms are delivery outputs, specified alongside features and verified before launch. An unmonitored invariant is an undetected incident.

Handover is the final gate

Production readiness ends with handover: runbooks the team has rehearsed, training completed, acceptance signed. The on-call rotation should inherit calm, not surprises — that is the standard we hold ourselves to on every programme.

Related articles