Backups that survive ransomware: immutability, isolation and restores you have actually tested

Ransomware operators go looking for the backups first. What it takes for a backup to survive: copies that cannot be changed, credentials the attacker cannot reach, and restores that have been rehearsed.

Almost every organisation has backups. Far fewer have backups that would still be there, intact and restorable, after an attacker has spent a week inside the network with administrator rights. That is the scenario worth designing for, because most serious ransomware today is operated by people rather than spread by worms, and those people know that the backup system is the one thing standing between them and a payment.

Why the backups are the first target

A human-operated ransomware attack is rarely a single event. There is an initial foothold, often a stolen password or an unpatched internet-facing service, followed by days or weeks of quiet work: mapping the network, escalating privileges, and finding the things that would allow the victim to recover without paying. Deleting Windows shadow copies, disabling backup jobs and wiping backup repositories are well-documented steps in that playbook. Encryption comes last, frequently timed for a weekend or a public holiday.

The environments that fare worst share a pattern. The backup server is joined to the same domain as everything else, administered with the same accounts, and reachable from the same network. Whoever holds domain administrator rights holds the backups too. The rule of thumb is blunt: if an administrator can delete a backup, so can anyone who has stolen that administrator’s account.

Snapshots and replicas are not backups

Three things are regularly mistaken for backups.

Snapshots on the same storage system are fast and useful for rolling back a bad patch, but they share fate with the storage they sit on and with the credentials that manage it. On a hyperconverged cluster, storage-level snapshots live on the same cluster as the data they protect.

Replication to a second site copies changes faithfully, including the encrypted ones. A replica is there for availability after a hardware or site failure. It only helps with ransomware if it keeps point-in-time history that the attacker cannot reach.

File sync services have the same limitation. Version history helps, but it sits under the same account the attacker may already control.

For this purpose, a backup is a point-in-time copy on separate storage, controlled by separate credentials, that cannot be modified or deleted before its retention period expires.

Immutability: copies that cannot be changed

There are several ways to make a copy immutable, and they can be combined.

Object storage with object lock writes data in a write-once, read-many mode. Implementations typically offer a governance mode, which suitably privileged users can override, and a compliance mode, which nobody can override until the retention period ends, including the account owner. For ransomware protection, the second is the one that matters.

A hardened backup repository is a dedicated Linux server that stores backup files with the operating system’s immutability protections set, has no remote administrative access beyond what the backup software strictly needs, and is not a member of any domain.

Offline media, meaning tape or removable disks that are physically disconnected once written, remains effective precisely because nothing on the network can reach it.

Two caveats apply. Immutability is only as strong as the management plane of the storage providing it: if an attacker can reach that console and shorten retention or delete the account, the protection is weaker than it looks. And immutable data cannot be cleaned up early, so capacity has to be planned for the full retention period.

Isolation: credentials, networks and the management plane

Immutability protects the data. Isolation protects the system that manages it.

  • Keep the backup infrastructure out of the production domain, with its own accounts and multi-factor authentication on every one of them.
  • Make the backup repositories reachable only from the backup servers, on the ports they need, with management interfaces on a separate network.
  • Prefer designs where the backup system pulls data from production, rather than production servers holding credentials that can write to, or delete from, the backup store.
  • Hold at least one copy under different administrative control, such as a different provider, account or set of people, so that no single compromised identity can reach every copy.
  • Alert on deleted backups, disabled jobs and changed retention settings. They are rare in normal operation and very informative when they happen.

Retention that outlasts the attacker

Because attackers are often inside the network well before anything is encrypted, the most recent backup may already contain their tools, or data they have tampered with. Retention needs to reach back far enough to find a clean restore point, which usually means keeping a tail of weekly and monthly copies rather than only the last couple of weeks of dailies.

It also means treating any restored system with suspicion. A server restored from a backup taken after the initial compromise brings the attacker’s persistence back with it. Restores after an incident go into an isolated network first, and are checked before being reconnected.

Restores you have actually tested

A backup job that reports success has proved that data was written. It has not proved that anything can be restored from it. Testing has levels, and each one catches different failures:

  • single-file restores, done routinely, which prove the catalogue and the credentials work
  • full virtual machine restores, done on a schedule, which prove a server will boot
  • a complete application restored into an isolated network, which proves the dependencies are understood
  • a whole-environment exercise, which proves the order of operations and the people

Order matters more than most plans admit. Identity comes first, meaning the domain controllers and DNS, then databases, then the applications that depend on them. An application server restored before the directory it authenticates against is a very expensive way to display an error.

Time matters too, and physics is not negotiable. Twenty terabytes pulled across a one-gigabit link takes around 44 hours at full line rate, before any overhead. If the recovery time the business expects is shorter than that, the design needs a local copy, a faster path, or both.

Finally, decide where the break-glass credentials for the backup system live. If they are in a password manager that runs on the infrastructure being restored, the plan has a hole in it.

A practical starting point

The 3-2-1 rule of three copies, on two kinds of media, with one offsite, is still a sound baseline. It is often extended to 3-2-1-1-0, adding one copy that is offline or immutable and zero errors in restore verification. The ASD’s Essential Eight lists regular backups as one of its eight strategies; its maturity model expects restoration to be tested, and its higher levels progressively restrict who can access, modify and delete backups.

Four questions are worth answering this month, honestly and preferably by testing rather than assuming:

  • Could someone with domain administrator rights delete or encrypt our backups?
  • Is at least one copy immutable or offline?
  • When did we last restore an entire server, and how long did it take?
  • Do we know the restore order, and where the break-glass credentials are?

Where Bizix fits

When Bizix builds a private cloud, backups are part of the hardening phase of the deployment runbook, alongside telemetry and monitoring, rather than something added after the first VM goes live. The environments we operate are covered by 24/7 incident response from the team that designed them, which is who you want on the phone on the morning these questions stop being hypothetical.