الدعم 24/7/365 · المكتب 9–17 بتوقيت الجبال، الاثنين–الجمعة +1 (208) 391-7176 team@alcousa.org

Backups are not a recovery plan

Almost every organization we assess has backups. A much smaller number have ever restored one. The gap between those two facts is where most of the damage in a serious incident actually happens.

There is a conversation we have had enough times to recognize the shape of it early. An organization has had an incident — ransomware, a failed migration, a deleted database, a hardware failure that took the array with it. We ask about backups. They have backups. The backup software reports success. Nobody looks worried yet.

Then we start restoring, and the questions begin. How far back does this go. Why is the database from Tuesday but the file share from Sunday. Why does the restore of a 4TB volume take eleven hours over this link. Why is the application server in the backup set but not the license server it depends on. Why is the encryption key for the backups stored on the domain controller we are trying to rebuild.

None of those are backup failures. Every one of them is a recovery failure, and they are only discovered during a recovery.

The two numbers that define your plan

Everything else in this article follows from two numbers, and if your organization cannot state them, that is the finding.

Recovery Point Objective is how much data you can afford to lose, expressed as time. An RPO of four hours means that in the worst case, you accept losing up to four hours of work. It is set by how much re-keying, re-processing, or lost transaction value your business can absorb.

Recovery Time Objective is how long you can afford to be down. An RTO of two hours means that from the moment of failure, the service is expected to be usable again within two hours. It is set by what the downtime costs, in revenue, in contractual penalties, and in the harder-to-price currency of customer patience.

These are business decisions wearing technical clothing. They should be set by whoever owns the consequence, not by whoever owns the servers. And they will differ enormously between systems in the same organization — a manufacturing scheduling system might need an RTO of thirty minutes while the internal wiki can be down for a week without anyone noticing.

Realistic recovery time by design, in hours

34Offsite tape or cold archive
11Nightly cloud image
4Replicated with daily snapshot
1Warm standby, tested
Indicative ranges for a mid-sized production workload, drawn from ALCO recovery testing. Actual times depend on data volume, link capacity and application complexity.

The chart is the whole argument. The same data, protected four different ways, produces recovery times that differ by more than thirty hours. The backup is not what determines that. The architecture around the backup is.

The failure modes nobody plans for

In our experience these are the five that actually bite, roughly in order of how often we see them.

  • The dependency you did not back up. The application is restored. It will not start, because it authenticates against a service that was out of scope, or reads a config from a share nobody catalogd, or needs a license server that lived on a VM somebody decommissioned.
  • The restore that is slower than the business can bear. Backups optimize for write speed and storage cost. Restores are the read path, and nobody tested it. Eleven hours of restore against a two-hour RTO is a plan that does not exist.
  • The backup that was in scope of the incident. Ransomware operators look for backup infrastructure specifically, because encrypted backups are what turns an incident into a payment. If the backup server is domain-joined, reachable from a workstation, and using an account with broad rights, it is not a backup — it is a second copy in the blast radius.
  • The silent failure. The job reports success because it completed. Nobody noticed it stopped including a directory eight months ago when a path changed.
  • The key nobody can find. The backups are encrypted, correctly, and the key material is stored inside the environment being recovered.
94%Of organizations hit by ransomware said attackers tried to compromise backups
57%Of those attempts succeeded
2xMedian ransom demand where backups were compromised
Sophos, State of Ransomware 2024, survey of 5,000 IT and cybersecurity leaders across 14 countries. Figures relate to organizations that experienced a ransomware attack in the preceding year.

That middle number is the one that should change your architecture. More than half of deliberate attempts on backup infrastructure succeed. The design assumption cannot be that your backups will survive an incident by default. It has to be that they are a target, and are protected accordingly.

What actually protects a backup

The principle is separation, applied consistently. An attacker who owns your production environment should not thereby own your recovery capability.

Immutability means the backup cannot be altered or deleted for a defined retention window, by anybody, including an administrator with valid credentials. Object-lock storage and hardened appliances both provide this. It is the single highest-value control in this list, because it is the one that holds even when credentials are lost.

Separate identity means the backup system does not authenticate against the same directory as production. If your domain is compromised, a domain-joined backup server is compromised with it.

Network separation means the backup infrastructure is not reachable from general workstation networks, and initiates its connections outbound rather than accepting them inbound.

Offline or logically air-gapped copies mean at least one copy is not continuously connected. The traditional 3-2-1 formulation — three copies, two media types, one offsite — remains sound, and the modern extension 3-2-1-1-0 adds one immutable or offline copy and zero errors on verification.

Key custody outside the environment means exactly what it says. If recovering the environment is a prerequisite for accessing the keys that recover the environment, you have built a circular dependency into your worst day.

The test is the plan

A recovery plan that has not been executed is a hypothesis. This is not a rhetorical flourish; it is the operational reality, and it is why we treat restore testing as a deliverable rather than an internal chore.

A real test restores to isolated infrastructure, brings the application up, and has somebody who uses the system daily confirm it works. Not that the files are present — that the thing does its job. Timing is recorded against the RTO. Data currency is recorded against the RPO. Anything that had to be improvised gets written into the runbook, because the improvisation is the finding.

Do this quarterly for tier-one systems and annually for the rest. Every test we have run for a new client in the last two years has surfaced at least one dependency that was not in the backup scope. Not most. Every one.

Where this lands

The uncomfortable version of this article is short. Backup software is a commodity and it mostly works. Recovery is an engineering discipline and it mostly is not practised. The organizations that recover quickly are not the ones that spent the most on backup licensing; they are the ones that decided their RTO in advance, designed backwards from it, put their backups outside the blast radius, and proved the whole thing worked while nothing was on fire.

If you cannot state your RTO, or you can state it but have never measured against it, that is the place to start. It costs a day. It is the cheapest insurance in this entire discipline, and it is the one almost nobody buys until after they have needed it.

المزيد من ALCO

Least privilege without breaking everything

Removing local administrator rights is one of the highest-value security changes available and one of the most frequentl

اقرأ المزيد

What an hour of downtime actually costs

The published averages are enormous and nearly useless, because they average across organizations nothing like yours. He

اقرأ المزيد

Infrastructure as code for organizations that are not software companies

The practice is usually explained by software companies, to software companies, with examples from software companies. T

اقرأ المزيد

هل هذا مناسب لك؟

نُقيّم ملاءمة كل مشروع قبل تسعيره. أي محادثة تقنية عن بنيتك التحتية، لا مكالمة مبيعات — وإجابة صريحة إن لم نكن الشركة المناسبة.

دعنا ننظر فيما تُشغّله.

محادثة تحديد نطاق مع مهندس أول. سنخبرك بما سنغيّره، وبكم سيكلّف، وبما إذا كنا الشركة المناسبة لذلك.