A practical structure for an enterprise tasked with moving from standard on-prem (Rubrik/Cohesity) + cloud/SaaS (AWS/Azure native backup) protection to a ransomware-resilient, air-gapped, immutable recovery posture.
1. Program Foundation & Governance
- Executive sponsorship and charter — Secure a named executive sponsor (CISO or CIO level) and a written charter defining scope, budget authority, and success criteria. Without this, vaulting and immutability initiatives stall at the funding stage because they compete with feature work that shows faster ROI.
- Cross-functional steering committee — Stand up a recurring working group with security, infrastructure, application owners, legal/compliance, and risk management. Ransomware recovery is not a backup team problem alone; decisions about RTO/RPO tradeoffs and vault access require buy-in from people outside the backup org.
- Data classification and criticality tiering — Inventory systems and classify them (Tier 0/1/2/3) based on business impact, regulatory exposure, and interdependency. This tiering drives everything downstream: which workloads get vaulted, how often, and how fast they must recover.
- RTO/RPO targets by tier — Define recovery time and recovery point objectives per tier, validated with business unit owners rather than assumed by IT. A common failure mode is building an expensive vault for Tier 0 systems while never formally agreeing what “acceptable downtime” means to the business.
2. Risk Assessment & Current State
- Backup environment audit — Document every backup domain (Rubrik/Cohesity on-prem, AWS Backup, Azure Backup, any SaaS-specific tools like for M365 or Salesforce) including retention, encryption status, and existing immutability features. Most environments have partial immutability already enabled (e.g., Rubrik’s native ransomware protection) that isn’t being fully leveraged or tested.
- Attack surface mapping for backup infrastructure — Identify every credential, service account, and network path that could reach backup consoles or repositories. Backup infrastructure is now a primary ransomware target itself; attackers routinely go after the backup admin console before triggering encryption.
- Gap analysis against a framework — Map current state against NIST CSF 2.0 or NIST SP 800-209 (Security Guidelines for Storage Infrastructure) and document gaps in isolation, immutability, and recovery testing. This gives the program a defensible, auditable baseline rather than an ad hoc wish list.
- Tabletop exercise (baseline) — Run an initial ransomware tabletop with IT, security, and leadership before building anything new, to expose assumptions and communication gaps. Teams frequently discover during the first exercise that nobody agrees on who has authority to declare a disaster.
3. Architecture: Isolation & Air-Gapping
- Logical air gap via immutable, isolated vault — Design a vault architecture (VaaS or private managed) that is network-isolated from production Active Directory and management planes, not just a separate storage target. True resilience requires the vault to survive a full domain compromise, meaning it cannot share credentials, DNS, or trust relationships with production.
- Vault deployment model decision (VaaS vs. private managed) — Evaluate managed vault-as-a-service offerings (e.g., Rubrik’s cyber recovery vault, Cohesity FortKnox, Commvault Cloud Air Gap Protect) against a self-built private managed vault in a separate data center or cloud account. VaaS reduces operational burden and speeds deployment; a private vault gives more control over data sovereignty and long-term cost at scale.
- Immutability enforcement (WORM) — Implement write-once-read-many locking at the storage layer (S3 Object Lock, Azure Immutable Blob Storage, or vendor-native immutability) with retention locks that cannot be shortened even by an administrator. This is the technical control that actually defeats ransomware and insider deletion, not just versioning or soft-delete.
- One-way data replication into the vault — Architect data flow so the vault only ever pulls or receives data, with no standing inbound path back to production that an attacker could exploit. A vault that can be written to by a compromised jump host isn’t air-gapped, it’s just a secondary target.
- Time-delayed or scheduled connectivity (“data bunker” pattern) — Where possible, limit vault connectivity to short, scheduled windows rather than an always-on link. This mirrors classic tape air-gap logic in a modern cloud context and shrinks the attack window significantly.
4. Cloud & SaaS-Specific Protection
- Native cloud backup hardening — Layer AWS Backup Vault Lock and Azure Backup’s immutable vault settings on top of existing AWS/Azure backup jobs, since native tools alone are often left in default (mutable) configurations. This closes the gap where cloud backups look protected but are actually deletable by anyone with the right IAM role.
- Cross-account/cross-tenant isolation for cloud backups — Store immutable copies in a separate AWS account or Azure subscription with its own IAM boundary, ideally with break-glass-only access. This prevents a compromised primary account (via stolen root or admin credentials) from reaching or deleting the backup copies.
- SaaS data protection coverage — Extend the program explicitly to SaaS platforms (M365, Salesforce, Google Workspace, Slack, etc.) using third-party SaaS backup tools, since most SaaS vendors’ native retention is not equivalent to a real backup. This is a frequently missed gap since teams assume “the SaaS vendor handles it.”
5. Ransomware Detection & Recovery Readiness
- Anomaly detection on backup data — Deploy ransomware detection capabilities that scan backup snapshots for encryption signatures, entropy changes, or mass file modification patterns before restore. Catching the infection inside the backup data itself lets you pick a known-clean recovery point instead of guessing.
- Clean room / isolated recovery environment — Build an isolated network segment (on-prem or cloud) where systems can be restored, scanned, and validated before reconnecting to production. Restoring directly into a still-compromised network risks reinfecting the very systems you just recovered.
- Recovery runbooks per tier — Document step-by-step recovery procedures for each criticality tier, including who executes each step, in what order, and with what validation checkpoints. Generic “restore from backup” documentation fails during a real incident when decision fatigue and time pressure are high.
- Credential and identity recovery plan — Explicitly plan for rebuilding or restoring Active Directory/Entra ID and PKI, since these are common ransomware targets and are prerequisites for restoring everything else. Programs that focus only on data recovery often stall because they can’t authenticate into anything to restore it.
6. Testing & Validation
- Scheduled immutable recovery testing — Perform full recovery tests from the vault on a defined cadence (quarterly at minimum for Tier 0/1) rather than testing only backup jobs. Backup success rates mean little if a full-scale recovery has never actually been proven end to end.
- Isolated recovery exercises simulating total compromise — Periodically simulate a scenario where production, AD, and primary backup infrastructure are all assumed lost, forcing recovery entirely from the vault. This is the real test of air-gap architecture, not routine file-level restores.
- Metrics and reporting cadence — Track and report recovery test success rate, actual RTO achieved versus target, and time-to-detect for anomalies, reviewed with the steering committee. This turns the program from a one-time build into an ongoing, auditable capability rather than a project that quietly goes stale.
7. Governance, Compliance & Sustainability
- Access control and least privilege for vault/backup admin — Implement role-based access, MFA, and just-in-time privileged access for anyone touching backup or vault consoles, with all changes logged and alerted. Backup admin accounts are frequently over-privileged and under-monitored relative to their actual blast radius.
- Regulatory and cyber insurance alignment — Map the program to relevant requirements (SEC cyber disclosure rules, industry-specific regulations, cyber insurance underwriting questionnaires). Many cyber insurance policies now require proof of immutable, air-gapped backups as a condition of coverage or a factor in premium.
- Documentation and audit trail — Maintain current architecture diagrams, runbooks, test results, and policy documents in a location accessible during an actual incident (not solely on the network that might be down). This is both an operational necessity and typically an audit requirement.
- Continuous improvement loop — Build a formal process to update the program as threats evolve, new tooling emerges, or the environment changes, with an annual full program review. Ransomware tactics evolve faster than most backup architectures; a program that isn’t revisited annually drifts out of relevance.
8. Identity Resilience (Often the #1 Failure Point)
- Active Directory forest recovery strategy — Define a dedicated AD recovery architecture including system state backups, forest recovery procedures, offline-secured administrative credentials, and documented rebuild sequences. Backups are useless if administrators cannot authenticate, and most ransomware recoveries stall because teams pursue application recovery before re-establishing trusted identity services.
- Entra ID / cloud identity recovery — Protect Entra ID configuration, Conditional Access policies, MFA settings, administrative roles, application registrations, and privileged groups. Many organizations run cloud workloads fully dependent on Entra ID with no tested recovery strategy for the identity layer itself.
- Privileged Access Workstations (PAWs) — Separate recovery administration from daily administration using dedicated hardened workstations never used for email or web browsing. If the administrator desktop is compromised, vault access cannot be trusted.
- Break-glass accounts — Create, test, and regularly validate emergency administrative accounts stored under strict controls. Recovery activities should never depend entirely on normal production identity systems remaining operational.
9. Recovery Prioritization & Dependency Mapping
- Application dependency mapping — Document application-to-database, application-to-identity, and application-to-network dependencies. Teams frequently discover mid-incident that a restored application can’t function because a supporting service wasn’t recovered first.
- Minimum Business Viable Operations (MBVO) — Define what the organization needs to operate at reduced but acceptable capacity. Most businesses don’t need full recovery on day one; recovering 20% of systems can restore 80% of critical business function.
- Recovery sequencing framework — Develop a restoration order aligning infrastructure, identity, databases, applications, integrations, and end-user services. Recovery efforts fail when teams restore systems independently without coordination.
- Business service recovery catalog — Shift thinking from recovering servers to recovering business services. Executives care about payroll, patient care, ERP, manufacturing, and customer transactions, not individual VM restore status.
10. Secure Recovery Environment
- Cyber recovery clean room — Establish a permanently maintained recovery environment isolated from production for validation and forensic review. Waiting until after an attack to build one creates significant delays when time matters most.
- Golden image repository — Maintain validated server, VM, workstation, and container images in immutable storage. Clean rebuilds are often faster and safer than in-place remediation of a compromised system.
- Secure CI/CD recovery — Protect source code repositories, container registries, image repositories, and deployment automation for pipeline-built applications. Applications cannot be rebuilt without them.
- Infrastructure-as-code protection — Protect Terraform, CloudFormation, ARM templates, Bicep templates, Ansible playbooks, and automation repositories. Recovery speed improves substantially when infrastructure can be redeployed from code rather than rebuilt by hand.
11. Security Operations Integration
- Security incident integration — Integrate backup and recovery teams directly into the incident response process. Recovery activities should begin while containment and investigation are still underway, not after security operations conclude.
- Backup security monitoring — Forward backup platform logs to the SIEM so retention policy changes, deletion attempts, vault access, failed MFA events, and privilege escalations generate security alerts.
- Threat hunting for backup infrastructure — Conduct periodic reviews focused specifically on backup systems. Modern ransomware groups often spend weeks targeting backup infrastructure before launching encryption.
- Insider threat controls — Protect against malicious administrators through separation of duties, approval workflows, immutable retention controls, and extensive audit logging.
12. Recovery Assurance and Validation
- Recoverability validation program — Move beyond “backup completed successfully” to verifying operating systems boot, databases mount, applications function, and business transactions can be performed.
- Automated test recovery — Use Rubrik, Cohesity, cloud-native, or third-party recovery verification capabilities to run routine automated testing at scale rather than relying on manual spot checks.
- Recovery certification process — Formally certify critical systems as recoverable through periodic testing, tracking systems that fail exercises the same way security vulnerabilities are tracked.
- Recovery scoring dashboard — Build executive-level reporting on recovery readiness score, percentage of workloads tested, vault coverage, immutable coverage, recovery success rates, and RTO/RPO compliance. This gives leadership a measurable cyber resilience maturity indicator instead of a status update.
13. Data Integrity & Recovery Point Validation
- Known-good recovery point identification — Develop procedures to identify the last uncompromised recovery point, since some ransomware campaigns dwell in an environment for months before activation.
- Malware scanning of recovery copies — Scan vault-hosted recovery points before restoration to prevent restoring an already-infected system.
- Data integrity validation — Validate database consistency, application functionality, and transactional integrity after restoration rather than relying solely on a successful recovery job status.
- Backup data retention analysis — Confirm recovery point retention extends beyond average attacker dwell time. A 30-day backup history may be insufficient if compromise goes undetected for 90 days.
14. Operational Resilience
- Recovery staffing model — Define primary, secondary, and tertiary recovery personnel, since incidents often occur during vacations, holidays, or simultaneous organizational disruptions.
- 24×7 escalation process — Establish documented escalation trees covering security, infrastructure, cloud, networking, applications, legal, and executive stakeholders.
- Vendor recovery support agreements — Confirm emergency support arrangements exist with Rubrik, Cohesity, AWS, Azure, security vendors, and key application vendors. Some organizations discover mid-crisis that their support contract is business-hours only.
- Recovery communications plan — Create out-of-band communications independent of production email and collaboration platforms, since Teams, Exchange, Slack, and VPN access can all go down simultaneously during a major incident.
15. Third-Party and Supply Chain Resilience
- Critical vendor dependency assessment — Identify business-critical third-party providers whose outage would impact recovery activities.
- MSP and MSSP resilience validation — Review the recovery and cyber recovery capabilities of managed service providers involved in backup or infrastructure administration.
- Software supply chain protection — Protect source code, artifacts, software repositories, and deployment pipelines against compromise.
16. Program Metrics & Maturity Model
- Cyber recovery maturity assessment — Implement a maturity model measured across governance, recovery readiness, identity resilience, vault protection, testing maturity, recovery orchestration, and operational preparedness.
- Key Risk Indicators (KRIs) — Track untested critical workloads, vault replication failures, expired credentials, backup policy exceptions, systems lacking immutable protection, and critical systems exceeding RPO targets.
- Executive reporting — Translate technical readiness into business risk language: how much data could be lost, how long recovery would take, which business functions are at risk, and what residual exposure remains.