Your website is down. What should you do in the first 15 minutes?
A calm route from the first alert to recovery. Confirm the impact, pause changes and avoid losing time to random restarts.

An alert arrives, a customer cannot complete an order and the team chat fills with competing explanations. The first few minutes decide whether you narrow the problem methodically or change five things at once. The goal is straightforward: establish the impact, restore the service and preserve evidence of what happened.
First, confirm what is actually broken
Open the service from another network and inspect the last successful checks. Is the whole website unavailable, or just login or a particular API? One failed attempt cannot tell you whether the fault is in the server, DNS or the route between them. Record the time, URL and exact failure.
Distinguish a connection failure from an HTTP response. A 503 means the server responded but the service is unavailable; a timeout provides different evidence. If the homepage works, test the step the customer needs to complete. A healthy homepage does not necessarily mean a healthy product.
Pause changes and choose one coordinator
Stop further deployments and agree who leads troubleshooting. A second person can track impact and handle communication. Even a small team benefits from one shared record: what you know, what you are trying and the result. Uncoordinated changes can hide the original cause.
Look for the last change, then the bottleneck
Compare the start of the incident with deployments, configuration changes, DNS updates and database migrations. Timing is a clue, not proof. Inspect application errors, database connections, disk space and resource usage. A restart without this evidence may help briefly while erasing a useful trail.
Consider a safe rollback before improvising a production fix. After a database migration, check that the previous application version still supports the new schema. Restoring service takes priority over a complete explanation, but it must still protect the data.
Communicate before you know the cause
State which functions are affected and when the next update will arrive. Avoid promising a repair time without evidence. After a change, check the service externally and watch several subsequent checks. One successful page load is insufficient grounds to close the incident.
- Confirm the outage and its scope.
- Pause changes and record each intervention.
- Restore service with a safe, verifiable action.
- Check the customer journey and announce recovery.
What to take away
Prepare this procedure before an incident and attach an accountable contact to each monitor. An alert saves time only when it leads to a clear response.


