Home›Guides›What to do if…

What to do if…

My server is down: what should I do?

A failed server is handled in this order: understand the failure, find out whether the data is still readable, choose the last clean restore point, restore, and only then decide whether this server should have had a ready-to-use standby. Restoring before you have identified the clean point can mean overwriting the only copy that is still good.

Updated October 20263 min read5 sources cited

Key points

  • Write down the time and the symptom before touching anything: this will be your starting point for choosing the right copy.
  • Several machines affected, or files renamed en masse: this is an attack, not a failure. Isolate and follow the ransomware guide.
  • Do not keep restarting a server whose disks are making noise: every boot can finish off a dying disk.
  • Restore from the last job that succeeded and predates the incident, after opening a test file from that point.
  • Time the return to service: that is your actual RTO.

1. Diagnose, without switching everything off at random

Write down the time and the symptom: no network at all, blue screen, clicking disks, an application that refuses to open, an encryption message.

  • Power, switch, cable. A server that is “down” is sometimes just a dead link. Do the other machines respond? Does the NAS respond?
  • A single service. The machine boots, the application does not. That is not the same delay, nor the same restore, as a dead disk.
  • Several machines at once, or files renamed en masse. Treat this as an attack, not a hardware failure: cut Internet access for the affected network, disconnect the affected machines without shutting them down and go to Ransomware has just struck. Do not restore onto a network that is still burning.

If the physical server smells of burning, or the disks are no longer audible, and you have no copy, stop powering it on again and again: every boot can make a dying disk worse. The backup copy becomes the priority.

2. Determine whether the data is intact

Three situations:

  • The system is dead, but the data disks still respond from another connection or a live CD. You can make an emergency copy to a healthy disk, then restore properly. This emergency copy is no reason to skip the off-site backup: it may be incomplete.
  • The files are there and open. A partial software or hardware failure. A repair may be enough. Back up the current state before attempting destructive repairs, if that state is still sound.
  • The files are unreadable, missing or encrypted. Production is no longer a valid source. Only an earlier backup is.

3. Identify the last restore point

In the backup console, take the last successful job, and check that it predates the incident. If the failure is corruption discovered today but which began a week ago, yesterday’s job is a poor candidate. Open a test file from that point before launching the full restore.

Find out where the encryption key is. Without it, the restore point exists but remains unreadable.

4. Restore

  • Files only if the system is sound and only a folder is missing.
  • Entire server if the system is dead: an image restored to equivalent hardware or to a virtual machine. This is faster than reinstalling by hand, provided the image has been tested at least once in the year.
  • Do not restore over a disk that may hold the only recent, unbacked-up data until that doubt has been cleared.

If several servers need restarting, follow the order of dependencies: directory and network first, then databases, then applications, then workstations. The ANSSI, France’s national cybersecurity agency, recommends defining this restore order in advance, taking into account dependencies and how critical each application is.

Time it. That figure is your actual RTO.

5. Consider a DRP if the server is critical

If the downtime has already cost too much, or if no replacement hardware is available, a DRP lets you restart now on a standby instance, from the chosen point, while the hardware is repaired. If this server fails often, or if management no longer accepts the delay, it must be added to the DRP or the BCP after the incident, in writing, not just in that evening’s conversation.

Degraded mode (paper, another tool) is triggered in parallel with steps 3 and 4, not afterwards.

After the incident: the debrief

Within the week, write down what took longer than expected, what was missing (password, key, contact, hardware) and what changes in the plan. If the cause is an attack, keep the traces and logs: file a complaint with the police in your country before reinstalling the machines, and notify any personal data breach to your country’s data protection authority (for example the APD in Belgium, the CNPD in Luxembourg or the CNIL in France) within 72 hours (GDPR, Article 33).

At WeDoBack

WeDoBack can restore the entire server, including the system, software and settings, or just the files. The copies are kept away from the failed server, encrypted, with the key held by the customer. With the DRP, servers restart on standby instances from the chosen version, without waiting to buy a new machine; activation is billed per day. Support can be reached on +33 9 72 50 78 28, from 9:00 to 13:00 and from 14:00 to 17:30 (Paris time). Outside these hours, monitoring may have raised an alert, but assisted restoration waits until opening time, unless other arrangements are set out in the contract.

Frequently asked questions

Should I shut the server down?

For a confirmed hardware failure (burning smell, clicking disks), yes: stop restarting it. If you suspect an attack, isolate it from the network rather than shutting it down: its memory may hold evidence useful to the investigation, as cybersecurity authorities point out, including the ANSSI in France.

How long does it take to restore a server?

It depends on the volume, the bandwidth, the method (files or full image) and whether replacement hardware is available. Without a tested system image, expect anywhere from half a day to two days for a physical server. A DRP lets you restart on a standby instance without waiting for hardware.

Should anyone other than the IT provider be informed?

If the outage results from an attack and personal data is affected, the breach must be notified to your country’s data protection authority within 72 hours (GDPR, Article 33). Also inform your insurer if it covers cyber risk, and file a complaint with the police in your country before reinstalling the machines.

Need help now?

Do not restore anything until you have identified a clean copy. We can guide you.

Call +33 9 72 50 78 28or write to us

Dealing with an incident right now?

Our teams help you identify the right copy and restore it, Monday to Friday, 9 am to 1 pm and 2 pm to 5:30 pm (Paris time).