Resources 8 min read

Backup, Recovery and Cyber Resilience

Why backup alone is not resilience. Immutability, recovery objectives, restoration testing, ransomware considerations and the difference between having copies and being able to recover.

An opened hard disk drive in black and white, showing the platter and the read head arm.

Every organization that has been through a serious ransomware event learned the same lesson, and most learned it late: having backups and being able to recover are different conditions. The first is a procurement outcome. The second has to be demonstrated, and until it has been demonstrated it is a hypothesis.

This article is about closing that gap: what makes backups survive a deliberate attack, what recovery objectives actually mean, and why restoration testing is the only evidence that counts.

Backup was designed for a different problem

Traditional backup answers accidental loss: hardware failure, deletion, corruption. Those events do not seek out the backups.

A capable intruder does. Modern ransomware operators locate backup infrastructure deliberately and destroy or encrypt it before triggering anything visible, because the backups are the only thing that makes the demand refusable. That changes the design requirement. A backup system reachable with the same administrative credentials as the systems it protects is not a recovery capability against this threat; it is another target on the same network.

Recovery objectives, and using them honestly

Two numbers frame recovery planning, and NIST's contingency planning guidance uses them throughout, expecting alternate arrangements to be configured in accordance with recovery time and recovery point objectives.

Recovery time is how long the business can tolerate a system being unavailable. Recovery point is how much data it can tolerate losing, measured backwards from the failure. The second is set by backup frequency; the first by the restoration process.

These are business decisions expressed in technical terms, and they are frequently written by technologists guessing on the business's behalf. The conversation is more productive framed concretely: how long can this department work without this system, and what happens on the day it is gone. The answers vary enormously by system, which is the point, because they justify spending differently.

The honest test of a stated objective is whether it has been met in a rehearsal. An organization claiming a four hour recovery time that has never restored the system in question is stating an aspiration.

Immutability and separation

The design property that matters against a deliberate attack is that a copy cannot be altered or deleted within its retention period, by anyone, including an administrator with valid credentials.

That is achieved through immutable storage where the platform enforces retention, through offline copies that are not continuously connected, or through a separate account or tenancy with its own authentication and no trust relationship with the production environment. The common thread is that compromising production does not confer the ability to destroy the copies.

The older discipline of keeping multiple copies, on different media, with one held elsewhere, remains sound. What the current threat adds is the requirement that at least one copy be beyond the reach of the credentials that run the environment, which is a stronger condition than being in a different building.

Restoration testing, the only evidence that counts

A backup job reporting success proves that a job ran. It does not prove the data is complete, that the media is readable, that the restoration procedure works, or that the result functions.

The failures found in testing are consistent. Backups that succeeded while silently excluding a database that was open. A restoration procedure that depends on a system which is itself down. Encryption keys stored only in the environment being restored. A restore that completes and produces an application that will not start because a dependency was not in scope. A recovery time that turns out to be four days rather than four hours because nobody had measured the data transfer.

Testing does not have to be a full scale exercise. Restoring a sample of files weekly and a full system quarterly, and recording how long the latter took, produces most of the assurance. The recorded duration is what converts a recovery time objective from a target into a measurement.

What resilience adds to recovery

Recovery restores what was lost. Resilience is the broader property of continuing to operate, and it includes the parts that are not technology.

The questions are practical. Which business processes must continue, and is there a manual alternative for a period. Who declares an outage and who communicates it. Where staff work if the primary environment is unavailable. How customers and suppliers are told. How the organization operates for a week if the systems do not return that day.

Technology teams rarely own these, which is why they are absent from otherwise thorough recovery plans. A plan that restores the systems into an organization that has not decided who talks to customers is half a plan.

Ransomware, specifically

Ransomware recovery differs from disaster recovery in several ways that matter.

The environment being restored into may still be compromised, so restoring immediately can reinstate the intruder. The point to restore from has to be before the initial access rather than before the encryption, which is why the investigation timeline discussed in incident response and cybersecurity recovery drives the recovery decision. Data may have been taken as well as encrypted, so restoring service does not end the incident or the disclosure obligation. And the scale is usually larger than any disaster scenario rehearsed, because it affects everything simultaneously rather than one site.

The practical preparations that pay are a clean build capability that does not depend on the compromised environment, credentials for the recovery path held outside it, and a documented order of restoration agreed with the business in advance.

Retention, and the question of how far back

Retention is usually set from a storage budget and occasionally from a regulation. It should also be set from a threat model, because the relevant question is how long an intruder might be present before anyone notices.

Dwell time of weeks or months is ordinary rather than exceptional. A retention scheme that keeps daily copies for a fortnight and nothing older gives an organization no clean restore point if the compromise began six weeks ago. Every available copy contains the intruder.

That argues for a tiered scheme: frequent copies with short retention for ordinary recovery, and less frequent copies held considerably longer for the case where the timeline turns out to be long. The longer tier is cheap, because it is infrequent, and it is the one that makes a recovery possible at all in the worst case.

Retention also runs the other way. Data kept beyond its purpose is data that can be disclosed, and a backup set holding personal information a decade after the operational system deleted it is a liability rather than an asset. The retention schedule and the deletion schedule have to be the same conversation, which is part of the documentation practice covered in cybersecurity policies and documentation.

What is actually protected

Backup scope is usually inherited rather than designed, and the gaps are predictable: data held in software as a service applications, which many organizations assume the vendor protects to a standard it does not; cloud resources created outside the managed process; endpoint data that never reached a server; and configuration, as distinct from data, without which a restored database is a database and not a service.

The configuration point deserves emphasis. Rebuilding a service requires the infrastructure definitions, the network configuration, the certificates and the application settings. Where infrastructure is defined as code and that code is itself backed up, this is largely solved. Where it is not, the recovery depends on somebody remembering.

Evidence and reporting

For audit purposes, the useful artifacts are the protection scope with its exclusions, the retention schedule, evidence of the last successful restoration test with its date and measured duration, and the recovery objectives agreed with the business.

The restoration test record is the one assessors increasingly ask for, because it is the only item on that list that cannot be produced by a system that has never been exercised. Maintaining it is part of the practice described in continuous security and compliance monitoring.

Running it as a service

Backup is the control most likely to be reported as healthy while being unfit, because its failure mode is silent and its verification is optional. Job success is monitored; recoverability usually is not.

LABUSA operates backup monitoring, restoration testing and recovery validation within LABUSA's managed backup and cybersecurity service, as part of the cycle described in the managed cybersecurity lifecycle. The protection of the backup infrastructure itself follows the practices in identity and access management, because that is the path an attacker takes to it.

Sources

About LABUSA

LAB Information Technology Incorporated (LABUSA) is a trusted provider of managed IT solutions, empowering organizations with secure, efficient, and scalable technologies. With expertise spanning cybersecurity, cloud services, enterprise software, and data management, LABUSA helps clients modernize operations, strengthen compliance, and optimize performance. Our customer-focused approach ensures tailored solutions that align with organizational goals while maintaining the highest standards of reliability and security. Headquartered in Houston, Texas, LABUSA serves government agencies, corporations, and nonprofits across the United States and internationally.