How an incident gets closed
The automated part of this system produces a colour. Everything customers actually value about a status page is the sentence next to it, and that sentence is a commercial communication written under pressure — which is exactly why it is drafted in advance and published by a person.
Key takeaways
- The status changes automatically; the note is drafted and published by a person.
- Three drafts exist in advance, so nobody writes prose during an incident.
- Updates on a fixed cadence even when there is nothing new, because silence reads as worse.
- A resolution says what was affected, for how long, and whether anything needs redoing.
- Publish uptime monthly per journey. A single site-wide figure hides the checkout outage.
Drafts written in advance
Nobody writes well at 14:08 with a broken checkout and a phone ringing. So the three notes that cover almost every incident are written calmly in advance, stored, and presented for one-tap publication with the affected journey filled in.
The three drafts
- Investigating. “We are aware that {journey} is not working correctly and are investigating. We will update this page within 30 minutes.”
- Identified. “We have identified the cause of the problem with {journey} and are working on a fix. We will update this page within 30 minutes.”
- Resolved. “{journey} is working normally again. The problem lasted from {start} to {end}. {impact}”
- Nothing else is templated. Anything more specific is written in the moment, because a specific claim made from a template is how a status page says something wrong.
The thirty-minute promise in the first two is load-bearing and it is a commitment the system then enforces: a reminder fires at twenty-five minutes so that the promise is kept even when the incident is absorbing everybody.
Updating when there is nothing to say
- App integration
- Machine learning
- Management
- Front-end & mobile
The resolution note
The one that matters commercially, and it has to answer three questions a customer actually has: what was affected, for how long, and is there anything they need to do.
The third is the one most status pages omit and the one that generates support contacts. “Orders placed between 14:06 and 14:31 may not have gone through; please check your order history or contact us” prevents a hundred emails. “The issue has been resolved” generates them.
Monthly numbers
- Machine learning
- Security & identity
- Management
- Analytics
Publishing four numbers instead of one is slightly more embarrassing and much more honest. The checkout figure is the one a customer cares about, and averaging it against three journeys that were fine produces a headline that is technically true and practically misleading.
What not to publish
Response times, error rates and internal component status all belong on an internal dashboard rather than a public page. A status page has one audience with one question: can I do the thing I came to do. Adding graphs invites a reader to interpret data they have no context for, and during an incident that produces speculation rather than patience.
Next: what all of this costs to run.
All posts