top of page

Search Results

Search this site

43 results found with an empty search

  • Veeam Clean Room Recovery Runbook

    The first time you run a clean room recovery should not be the day you need it. Yet that is exactly how most organizations are set up today. They have backups. They have immutability. Some even have a document titled "Cyber Recovery Runbook" that nobody has opened since the auditor asked for it. What they do not have is a practiced, orchestrated path from "we've been hit" to "we are validating clean workloads in an isolated environment right now." A clean room is not a place. It is a workflow. It is the repeatable process of taking an immutable copy of your data, standing it up in an isolated environment, proving it is clean, and promoting it back to production with an audit trail. If your clean room requires heroics, it will fail under the pressure of a real incident. If it is orchestrated, documented automatically, and drilled quarterly, recovery becomes a routine instead of an event. This guide covers the full picture with Veeam. The clean room you can build today with Veeam Backup & Replication (VBR) and Veeam Recovery Orchestrator (VRO) for the workloads that matter most: orchestrated, documented automatically, and drilled on a schedule. What a Clean Room Actually Needs to Do An Isolated Recovery Environment (IRE) gets defined a lot of different ways in this industry, usually in whatever way the vendor's product the answer. Here is a baseline definition of what an IRE and Clean Room is: Requirement What it means in practice Establish a minimum viable company A business impact assessment that identifies the applications the business cannot operate without. If you cannot name your top 20 workloads and their boot order, start there. Validate clean with SecOps integration Scanning restore points for malware with alerts forwarded to the security team. Recovery point selection is a security decision, not a backup decision. Test both disaster and cyber recovery An automated, repeatable testing process. The same clean room should exercise both scenarios, because in real life you rarely know which one you are in for the first few hours. An isolated environment Logically air-gapped with compute available. Isolation must be architectural, not procedural. A checklist that says "remember to disconnect the network" is not isolation. An identity safety net Identity recovery tested at both item level and at scale. Identity is where most clean room designs quietly fall apart. Ensure survivable data An isolated, immutable copy of the data. The foundation everything else stands on. Who Owns What: The Cyber Recovery RACI To say technology is half the battle in a cyber recovery would be an oversimplification. People and process and the level of preparation prior to the incident are the majority of the battle in a cyber recovery. There is a misconception that the data protection team will choose the clean restore point. The recovery point decision belongs to the security team, and the execution belongs to the data protection team. Here is that expanded into the full matrix. R is responsible (does the work), A is accountable (owns the outcome, one per row), C is consulted, I is informed. Activity Security / IR Data Protection IT / App Owners Executive Legal / Compliance Declare incident, activate IR plan R C C A I Bound the compromise window (forensic timeline) A / R C C I I Select the clean restore point A R C I I Maintain immutable copies, 4-eyes, platform hardening C A / R I I I Stand up the clean room and isolation C A / R R I I Run threat scans and validate restore points A R C I I Choose the path: recover, rebuild, decrypt, accept loss R C C A C Remediate workloads in staging A C R I I Approve promotion to production C R C A I Re-protect remediated systems I A / R C I I Preserve evidence, apply legal hold C R I I A Regulatory and disclosure obligations C C I A R Run quarterly drills, publish results C A / R R I I The path decision is accountable to the executive, not to IT or security, because recover-versus-rebuild-versus-accept-loss is a business risk trade, and the drill row is accountable to data protection because a resilience program needs one owner who is graded on it. Fill in names, not team labels, and revisit the matrix at every drill. A RACI with a vacant seat is a RACI that fails silently. Alignment with Regulations and Frameworks Auditors have stopped asking whether you have backups. They are asking whether you have proven you can recover, measured, and with evidence. When organizations have a high clean room recovery maturity they meet a growing list of regulatory requirements. Regulation / framework What it asks for What Veeam provides DORA (EU financial entities) Backup policies with restoration and recovery procedures, plus digital operational resilience testing, including scenario-based testing of ICT recovery (Articles 11-12, 24-26). Quarterly clean room drills are scenario-based recovery testing by definition. VRO plan reports document execution, tested RTOs, and results, generated automatically. NIS2 (EU essential and important entities) Risk management measures covering backup management, business continuity, and crisis management (Article 21). Survivable immutable copies plus a drilled, documented recovery path from them. The maturity ladder is a defensible implementation of "backup management and business continuity." NYDFS 23 NYCRR 500 (NY financial services) Incident response and BCDR plans, backups maintained isolated from network threats, and periodic testing of the ability to restore from those backups (Section 500.16, as amended). Vault and hardened repositories are the isolated backups. The drill checklist is the restore testing. The plan reports are the proof. HIPAA Security Rule (US healthcare) A contingency plan: data backup plan, disaster recovery plan, and testing and revision procedures (164.308(a)(7)), with proposed rulemaking pushing toward mandatory testing and defined restoration timeframes. Tested RTOs against the minimum viable company answer where the rule is heading, not just where it is. Identity recovery drills cover the systems PHI access depends on. PCI DSS 4.0 (cardholder data) An incident response plan that is tested at least annually, with defined roles and communication procedures. The first-24-hours sequence, drilled quarterly, exceeds the annual testing bar, and the 4-eyes approval trail documents who did what. ISO/IEC 27001:2022 Information backup (A.8.13), ICT readiness for business continuity (A.5.30), and maintaining security during disruption (A.5.29). The clean room is ICT readiness made testable: an isolated environment where continuity is exercised rather than asserted. NIST CSF 2.0 The Recover function: executed recovery plans, verified restoration integrity, and recovery communications. Threat Hunter and YARA validation before promotion is restoration integrity verification. The drill cadence is plan execution. The reports are the communications artifact. SEC cyber disclosure (US public companies) Material incident disclosure on a four-business-day clock, plus disclosure of risk management processes. A bounded compromise window and tested RTOs turn the disclosure from speculation into facts. The preserved forensic timeline and evidence retention support what gets filed. The regulator, the auditor, and the cyber insurer do not want your policy PDF. They want the drill report with dates, tested RTOs, and scan results, which the orchestrated tier generates as a byproduct of doing the work. NOTE: This section maps capabilities to requirement themes and is not legal or compliance advice. Regulations evolve and interpretations vary by jurisdiction and sector. Validate specific obligations with your compliance and legal teams. Recover, Rebuild, Decrypt, or Accept Data Loss In our industry we assume clean room recovery is the only viable solution from a cyber incident. It certainly is the most optimal, but every business faces 4 options during a cyber event. They need to quick validate if they can recover cleanly because they proactively did everything this paper discusses, or they have to move to less optimal options, whether that be rebuilding, decrypting, or accepting the data loss. Path When it is the right call What it costs 1. Clean Room Recovery A clean restore point exists within acceptable RPO. The default outcome when the tiers in this guide are working. Recover to staging, remediate (IOC removal, patching, hardening), promote. Any data written after the clean point is lost unless salvageable at file or database level. Remediation effort still applies. Fastest path back by far. 2. Rebuild OS, recover data only Dwell time is long or unbounded, OS-level persistence or rootkits are suspected, or compliance demands a known-clean operating system. Build fresh from a golden image, then restore only the data. Slowest clean path. Requires maintained golden images and application reinstall discipline. Veeam's file-level, disk-level, and application-item recovery decouple the data from the compromised OS, and the portable backup format restores that data onto fresh builds on any platform. 3. Decrypt No clean copy exists anywhere. This is the last resort, and its presence on this list is honest acknowledgment that some organizations arrive here. Coveware by Veeam brings incident negotiation expertise, decryptor validation, and compliance checks to a path nobody should walk alone. Decryptors are slow, frequently buggy, and often partial. Payment guarantees nothing and may be legally restricted. Walking this path is evidence the resilience program failed. The goal of every other page in this guide is to make this row irrelevant. 4. Accept data loss A clean point exists but it is old, and the delta is low-value, reproducible, or re-creatable from other systems of record. Rolling back and consciously forfeiting the gap can beat days of forensics on data that was not worth the recovery cost. The forfeited data is gone. This is a business decision with executive sign-off, not an IT decision. Document what was lost and why the trade was made, because the auditor will ask. CDP journal granularity shrinks this gap from days to minutes for the workloads it protects, which is often the difference that keeps this path off the table. Recovery Point: Dwell Time Changes Everything Dwell time, meaning the window between initial compromise and detection, has been collapsing. Attackers used to sit in environments for weeks. Increasingly they are detected in days or even hours, partly because encryption-first attacks announce themselves and partly because detection tooling has genuinely improved. The recovery implication is significant: if dwell time is short, your most recent recovery points are probably clean. Short dwell time: storage snapshots and replicas are in play If the security team can establish that the compromise began 36 hours ago, the storage snapshot from three days ago is a legitimate, and dramatically faster, recovery point. Veeam can orchestrate recovery directly from storage snapshots on supported arrays and can fail over VM replicas into an isolated clean room network through VRO. Replicas are already hydrated, sitting in native VM format, waiting to boot. When the forensic timeline supports it, this is your lowest RTO path by a wide margin. The clean room process stays the same. Boot into isolation, scan, validate, promote. Only the source changes. CDP takes this further. Veeam Continuous Data Protection replicates vSphere VMs with an RPO measured in seconds, and Universal CDP, new in V13, extends that to any Windows or Linux machine, physical, virtual, or cloud, replicated to vSphere in a ready-to-start state. The clean room value is not just the low RPO. It is the journal. Nightly backups force you to choose between yesterday and the day before. The CDP journal lets the security team scrub backward through short-term restore points and land minutes before the stamped detection time, instead of forfeiting a full day of data to get behind the compromise. For the workloads where the delta is the business, revenue systems, order books, patient records, that granularity is the difference between a rollback nobody notices and a rollback that makes the news. Same honesty as replicas and snapshots: CDP replicas are a speed and granularity play, not an immutable copy. And with seconds-level RPO, the newest journal points may contain the attack itself. The discipline is scrubbing to a validated point behind the compromise window, then scanning it like any other candidate before promotion. Longer or uncertain dwell time: immutable backups When the security team cannot bound the compromise window, or when the window predates your snapshot and replica retention, you fall back to backup copies. This is where restore point scanning earns its keep. Veeam Threat Hunter and custom YARA rules (YARA is a pattern-matching language security teams use to describe malware signatures) scan restore points inside the clean room so only validated data goes back to production. Worst case: the offsite copy Production is untrusted, on-site infrastructure is untrusted, and possibly the backup infrastructure itself is in scope. This is what your offsite immutable copy exists for, and it is where Veeam Vault anchors the design. Scenario Recovery source Why Dwell time bounded in hours to days CDP replicas (seconds-level journal), storage snapshots, VM replicas Already hydrated, native format, lowest RTO. CDP journal granularity lands minutes behind the compromise instead of a day. Valid when the forensic timeline confirms the point predates compromise. Dwell time uncertain or longer Immutable backup copies (hardened repository, object lock) Scanned with Threat Hunter and YARA in the clean room before promotion. Backup infrastructure in scope Veeam Vault (offsite) Immutable, encrypted, logically air-gapped, independent of on-site infrastructure. SOC and SOAR Integration for Better Decision Making The compromise window does not get established by the backup team squinting at job history. It comes from the security stack, and this is where Veeam's bidirectional SOC and SOAR integrations earn their place in the clean room design. Detection platforms including Splunk, CrowdStrike, Palo Alto Networks, Fortinet, and ServiceNow integrate to enrich detection and automate response. The flow runs both directions: Veeam's scanning results, entropy analysis, and indicators of compromise forward into the SIEM where the security team already lives, and the security stack pushes back into Veeam through the Incident API. When the EDR or SOAR platform detects compromise, it can mark the affected restore points as suspect and, when configured, trigger an immediate quick backup, stamping the moment of detection directly onto the backup timeline. That stamp is the starting point for the clean-point search. Instead of the security team handing the backup team a timestamp over a bridge call at 3 a.m., the timeline is already annotated when the clean room work begins. It’s important to note the recovery point decision belongs to the security team, and the execution belongs to the data protection team. Survivable Backup Data: Any backup copy job destination is a legitimate clean room source, as long as it is immutable. A Veeam Hardened Repository on-prem. An S3-compatible object storage bucket with Object Lock. Vault in the cloud. All of these work. Immutability is the non-negotiable. If the attacker can modify or delete your source data, nothing downstream of it matters. That said, Veeam Vault is emphasized as the offsite copy for three reasons: It checks every survivable-data box in one move. Immutable, encrypted, offsite, and logically air-gapped. No infrastructure for you to patch, no storage array for the attacker to find credentials for, no capacity planning exercise. It is the 3-2-1-1-0 offsite copy without the second data center. It is storage, not compute. Several competing architectures run their backup storage on a scale-out compute layer. That means the offsite copy carries a compute bill whether you are recovering or not, and recovery through those platforms means standing up standby clusters before you can even retrieve metadata. Vault is built on cost-effective object storage. Data sits cheap and boots fast when needed. It is where clean-point identification happens. Threat Hunter and YARA scanning run against restore points in Vault, so the same copy that survives the attack is the copy you validate before recovery. No shuffling data between a vaulting product and a scanning product. Source Immutable Clean room fit Veeam Vault Yes, by default Emphasized offsite anchor. Survivable, encrypted, zero infrastructure, scan-in-place. Veeam Hardened Repository Yes Excellent on-prem source. Standard Linux deployment for the clean room copy (see requirements). S3-compatible object storage with Object Lock Yes Excellent. Supports read-only concurrent access from the clean room, simplifying the isolation cycle. CDP replicas, storage snapshots, VM replicas No (speed play) Valid short-dwell recovery points with security sign-off. CDP and Universal CDP add seconds-level journal granularity. They complement the immutable copy. They do not replace it. Protecting the platform itself: 4-eyes authorization A hardened backup platform is a prerequisite though to a survivable backup. Attackers do not just encrypt production anymore. They go after the backup infrastructure first, and the fastest path is a stolen backup administrator credential. Immutability protects the data on disk. 4-eyes authorization protects the control plane. With 4-eyes enabled, destructive operations like deleting backups, removing repositories, or changing security settings require approval from a second backup administrator before they execute. Clean Room You Can Build Today No new licenses. No new appliances for VDP Foundation or Advanced customers. Every component below ships with the platform you already run. The recipe: an isolated host or small cluster with no routes to production, a recovery VBR server inside that island, and your immutable copy as the source. As of V13.1, that recovery VBR can attach an immutable object storage repository in read-only mode. Instant recovery boots workloads directly from backup storage into the isolated network. SureBackup and DataLabs automate the validation: boot the VM in a sandboxed network, confirm the OS starts, confirm services respond, run malware scans, run custom scripts. Threat Hunter and YARA rules scan the restore points before anything is trusted. This has been shipping for years, and plenty of organizations run a credible clean room on exactly this and nothing else. For agent-based workloads, physical Windows and Linux servers restore into the isolated environment as VMs or via instant disk publish. Orchestrated Clean Room with VRO This is the tier where the clean room stops being a runbook and becomes a rehearsed capability. The reference architecture comes from Veeam's Solutions Architect team, and the full best practices page lives at bp.veeam.com. This section summarizes the design. Scope note up front: this workflow covers vSphere and Hyper-V image backups. Everything it does not cover can be easily done with just VBR. The on-prem clean room is a small, isolated environment containing a backup repository, an embedded Veeam Backup & Replication (VBR) server running inside the VRO appliance, and enough compute to boot and test your minimum viable company. Backup copies from production land on the clean room repository, which is immutable. The embedded VBR inside VRO reads that copy, imports the backups, and all restore and testing activity happens in full isolation from production. NOTE: The repository must never be written to by production and the clean room simultaneously. Concurrent write access is unsupported and can corrupt data. Object storage accessed read-only is safe, and as of V13.1, immutable object storage repositories attach in read-only mode as a supported configuration (see below). This is why immutability on the repository is strongly recommended even inside the clean room design. It protects the data from corruption or modification even if someone fat-fingers the isolation schedule. Requirements worth knowing before you build Requirement Detail vCenter inventory connection VRO requires an initial connection to the production vCenter to gather inventory and build restore plans using vSphere tags. Connection to production VBR is optional, needed only for plans based on backup jobs. Workload support vSphere and Hyper-V image backups today. Agent and NAS backups are not supported in this workflow yet, though scripted workarounds exist in the community library. Hardened Repository deployment Hardened Repositories deployed via the Veeam Infrastructure Appliance ISO are not currently supported in this workflow. Use a standard Linux Hardened Repository for the clean room copy. Repository sizing Prefer multiple smaller repositories with shorter retention. Rescan time during import scales with content, and rescan time is dead time in your RTO. Threat Hunter in the clean room Enable it on the embedded VBR in the VRO server. Internet access (proxy supported) is required for activation and current signatures. Post-connection wait After connecting a repository to the embedded VBR, wait 10 minutes before running plans. The Orchestrator needs that window to collect restore point information from VBR. Read-only repository mode Version 13.1 turns the concurrency workaround into a product feature. A second backup server can now attach an immutable object storage repository in read-only mode and use it for restores, recoverability checks, and forensic or ransomware scans, without taking ownership away from production. Read that through the clean room lens: older clean room designs required careful choreography so production and the clean room never touched the repository at the same time. For immutable object storage, read-only mode eliminates that risk at the platform level. The clean room VBR mounts the copy, scans it, restores from it, and can never modify it, while production keeps writing uninterrupted. Two implementation details matter. First, select read-only mode when connecting the repository on the second server. Second, the credentials used to add the repository should themselves carry read-only permissions, because full-access credentials behind a read-only mount can produce unpredictable behavior. Defense in depth applies to the clean room too: the mode constrains the software, the credentials constrain everything else. And to be precise about scope, read-only mode governs what the secondary server will do. It does not isolate the storage network for you, and it does not replace segmentation or immutability. It removes the concurrency problem, which was the most fragile part of earlier designs. This also strengthens a scenario that used to be a recovery problem in itself: a compromised or unavailable production backup server. Historically the path to the data ran through it. With read-only attach, a standby administrative environment is a supported configuration rather than a workaround, which is exactly what a clean room is. Stronger isolation: the Fibre Channel variant For organizations that want the strongest isolation, the clean room repository can be backed by an external storage system using storage snapshots or storage replication over Fibre Channel instead of the network. A dedicated Linux server inside the clean room mounts the replicated LUN or snapshot, and the embedded VBR imports the backups from there. No network path between production and clean room exists at all, and there is zero concurrency risk. This is also where replicas and storage snapshots shine as recovery points for short-dwell scenarios: the data is already sitting on an array the clean room can reach, in a format that boots immediately. Validation inside the room This is where the Veeam approach separates from the build-it-yourself model. SureBackup and DataLabs have automated isolated recovery verification on-prem for years: boot the VM in a sandboxed network, confirm the OS starts, confirm services respond, run malware scans, run custom scripts. VRO extends that into full orchestrated plans with automatically generated documentation and RTO/RPO reporting. When the auditor, the regulator, or the board asks "can you recover, and can you prove it," the answer is a scheduled report, not a scramble. Cloud Clean Room: Vault Plus Instant Recovery to Azure Not everyone can afford, or wants to maintain, a physical clean room. DR to the cloud is becoming a more common design. Vault for immutable storage and clean-point identification, an isolated Azure VNet for live forensic validation, and Instant Recovery to Azure as the mechanism that makes it fast enough to meet real RTOs. No VRO required. VRO layers orchestration on top, covered below. Instant Recovery to Azure boots the VM live in Azure before the full migration to native Azure Managed Disks completes, similar to how instant restore to VMware works on-prem. It is not limited to native Azure workloads. Backups of VMware, Hyper-V, AHV, Proxmox, HPE Morpheus, EC2, Azure, GCP, and physical Windows and Linux servers can all instant-restore into Azure. Read that list again through a clean room lens: your recovery target no longer depends on your production hypervisor surviving the attack. The workflow Backup copies land in Vault. Immutable, encrypted, offsite. Even if every snapshot and every on-prem copy is destroyed, this copy is intact. The security team bounds the compromise window. Veeam's Indicators of Compromise mapping to the MITRE ATT&CK framework and threat scanning against restore points in Vault feed the clean-point decision. Threat Hunter and YARA scan the candidate restore point before anything boots. Instant Recovery boots the workload into an isolated VNet with no routes to production. The isolation is architectural. There is no procedural step where someone can forget to disconnect something. Validation runs, then promotion. The workload either migrates to native Managed Disks and takes over, or it fails validation and you iterate to an earlier point. Only validated data goes forward. Orchestrating it with VRO VRO orchestrates this end to end with Azure Recovery Locations. Add the clean room Azure subscription (or a non-peered VNet in an existing subscription) to VBR, define the recovery location in VRO with network mapping to the isolated VNet, subnet, and security group, then attach PowerShell plan steps for verification and automated teardown of Azure resources after testing. Members of the Veeam community have documented this pattern in detail, including running full restore-verify-teardown cycles of a 250 GB VM a dozen times for pennies of Azure spend, because the resources exist only for the duration of the test. Operational points that save you during a real event Pre-deploy Helper Appliance Templates. Veeam publishes OS-specific helper templates to the Azure Compute Gallery in your target region, which meaningfully cuts time-to-first-boot. If your first instant restore happens during a live incident and the templates are not staged, you are donating time to the attacker. Make this a quarterly drill item. Use existing resource groups, NSGs, and storage accounts for mass restores. Azure API throttling is real, and it will bite you exactly when you are restoring 200 VMs at 2 a.m. The isolated recovery VNet is not optional. Booting a potentially compromised VM into a production VNet to "verify it works first" always ends the same way: it makes things worse. Build the isolated network. It takes 15 minutes. Use it every time. The cost advantage with Veeam Recovering from Vault means booting directly from Blob-backed storage with minimal compute and migrating in the background. The appliance-centric alternative means either restoring through a compute-heavy intermediary or landing in dedicated bare-metal cloud infrastructure like AVS or NC2 at premium pricing. The cloud bill during recovery is part of your total cost of protection whether the vendor quote mentions it or not. Identity: The Critical Path Nobody Drills Here is the uncomfortable truth about most clean room tests: they boot application VMs, watch the OS load, and declare victory. Then the real incident hits and nobody can log into anything, because identity was never part of the plan. Identity is the critical path to recovery. Applications do not work without authentication, and modern environments run overwhelmingly on non-human identities: service accounts, service principals, OAuth configurations, machine identities. If Active Directory or Entra ID is compromised or unavailable, your beautifully validated application VMs are furniture. A clean room design has to answer for both directions. Active Directory Entra ID Instant item-level recovery of individual AD objects directly from image-level backups, without restoring entire domain controllers, for the surgical cases. Object recovery at scale for users, groups, and directory roles without a full tenant rollback. Forest-level recovery that performs the full VM restore and runs post-restore activities to bring AD up cleanly, for the catastrophic cases. Side-by-side comparison of property values between a restore point and production, so you can see exactly what the attacker changed. Threat detection at the AD server layer, watching for the tooling and behaviors attackers use against domain controllers specifically, because AD is where attackers go first. Recovery of conditional access policies, service principals, and enterprise app settings, which is the part everyone forgets until the apps will not authenticate. Drill identity recovery both ways: item-level and at scale. It is called a safety net for a reason. It is the thing that catches everything else. Returning to Production The clean room is the middle of the story, not the end. The road back to production is where organizations get sloppy, because by this point everyone is exhausted and the pressure to declare victory is enormous. Four disciplines close it out properly. Last-mile validation Scan the remediated workloads one final time before promotion, with current threat signatures, not the ones from the day the incident started. Threat intelligence moves during a multi-day recovery, and the variant that hit you may have known indicators by day three that did not exist on day one. The final Threat Hunter pass is cheap insurance against promoting something the first scan could not see. Veeam has a named feature for exactly this: Secure Restore, which builds the Threat Hunter, antivirus, or YARA scan into the restore workflow itself, so the workload is verified as part of the restore rather than trusted after it. Re-protect immediately The first job after promotion is a backup job. The remediated, validated state of every recovered workload goes straight to the immutable copy as the new baseline. If anything resurfaces, you restore to the remediated state, not back to the pre-incident state that started this whole exercise. Skipping this step means your best restore point is still the compromised one. 4-eyes on the promotion Promotion touches production, and production-touching operations during an incident are exactly when mistakes and malice hide best. 4-eyes authorization puts a second administrator on the destructive and irreversible actions, and the approval trail doubles as the change record your post-incident review and your insurer will both want. Preserve the evidence The infected restore points are evidence. Law enforcement and insurers may prohibit deleting them until investigations close, and your own post-incident review needs them. Put the compromised restore points and the forensic timeline under extended retention on the immutable copy, separate from the operational backup chain. Plan the storage for this in advance, because "we deleted the evidence to free up capacity" is a sentence you do not want to say to a regulator. Clean Room Maturity Self-Assessment Grade yourself honestly across seven dimensions and four levels. Dimension Level 1: Ad hoc Level 2: Foundational Level 3: Practiced Level 4: Proven Survivable data Backups exist. Immutability unverified. Immutable copy (hardened repo or object lock). Offsite copy exists. Vault or equivalent offsite anchor. 3-2-1-1-0 verified, dedupe caveats addressed. Read-only standby attach (V13.1). Evidence retention tier planned and funded. Recovery point selection Latest backup, always. Security team consulted informally. Dwell-time-driven: CDP journal, snapshots, replicas, or backups by compromise window. Incident API stamps the timeline. Selection rehearsed with security in drills. Isolation None. Restores land in production. Manual isolated network, procedural controls. Architectural isolation: isolated host or non-peered VNet, built and reusable. Isolation is standing infrastructure, exercised quarterly, on-prem and cloud. Validation Boot it and hope. Antivirus scan after restore. Threat Hunter and YARA on candidates. Secure Restore in the workflow. Last-mile scan with current signatures gates every promotion. Results feed the SIEM. Identity recovery Not part of the plan. AD backups exist. Item-level restore tested once. AD forest and Entra ID recovery both tested. Item-level and at-scale. Identity boots first in every drill. Conditional access and service principals covered. Orchestration and proof Tribal knowledge. Written runbook, manually executed VRO plans for the minimum viable company Documentation auto-generated. Plan reports feed audits, insurers, and regulators without manual assembly. Governance and drills No drills. No named owners. Annual test. RACI drafted. Quarterly drills. RACI filled with names. 4-eyes enabled. Drills include workloads and path decisions. Gaps reported to executives honestly. Where the market actually sits: most organizations self-assess at Level 3 and drill at Level 1. The gap between what the runbook says and what the team has done with their hands is the single best predictor of how a real incident goes. Everything in this guide exists to move you one level to the right, one dimension at a time, starting with your lowest score. Where the Assemble-It-Yourself Approach Falls Short There is a school of clean room design built around scale-out backup appliances. The recipe: buy or stage a standby cluster, assemble a "digital jump bag" of ISOs and golden images in a file share, manually create isolated VLANs and virtual IPs on the cluster, retrieve metadata from the cloud vault to the standby cluster when disaster strikes, then walk through recovery via a lengthy deployment guide. These guides are thorough. They are also 60-plus pages of procedures a human executes under maximum stress, and every manual step is a place the recovery can stall. Four structural problems with that model: Structural problem The appliance model The Veeam model Idle infrastructure A standby cluster is capital expenditure waiting for a bad day. Software you already own, data in a portable format, and either modest on-prem compute or cloud compute that exists only during recovery and testing. Storage coupled to compute The offsite copy is expensive at rest and the recovery path runs through more of the vendor's own infrastructure. Vault decouples them. Cheap at rest, fast in recovery. Procedural isolation Manual VLAN creation, manual VIP assignment, manual firewall rules, verified by a human pinging things. Works until someone skips a step. Architectural isolation: a non-peered VNet or a scheduled network window managed by scripts and orchestration. Does not depend on memory. Testing as an event When recovery is a manual runbook, testing is a project, so it happens annually if at all. Orchestrated plans make testing a scheduled job with automatically generated documentation. Recovery muscle is built through repetition. Data portability deserves more attention than it gets. Veeam backups are portable files. The same backup of a VMware VM can restore to VMware, Hyper-V, Nutanix AHV, native Azure, AWS, Google Cloud, or bare metal. Your clean room is not locked to the platform the data was born on. When the appliance model is your foundation, your recovery options are whatever the appliance supports, wherever the appliance lives. The Quarterly Drill Checklist # Drill item 1 Confirm your minimum viable company list is current, with boot order and dependencies. 2 Run a full clean room cycle from at least two sources: your fastest (replica or storage snapshot) and your survivable (Vault or hardened repository). 3 Recover at least one Tier 1 workload (an agent-based physical server or cloud-native machine) alongside the orchestrated Tier 2 plan, so the manual path stays practiced too. 4 Verify Helper Appliance Templates are staged in every Azure region you would recover into. 5 Run Threat Hunter and a YARA rule against a real restore point and confirm alerts reach the security team. 6 Recover identity both ways: one AD object item-level, and one forest or tenant-scale exercise. 7 Run the Security & Compliance Analyzer against the backup infrastructure and remediate any configuration drift it finds, so the platform itself stays hardened between incidents. 8 Time everything. Compare against your stated RTOs. Report the gap honestly. Summary The organizations that recover fastest treat recovery as a first-class engineering discipline. They do not buy a clean room. They practice one. The Veeam version of this discipline is simple to describe even if the engineering behind it is not. Survivable data in Vault or any immutable copy destination. The fastest safe recovery point, which is increasingly a replica or storage snapshot as dwell times collapse. Architectural isolation on-prem or in Azure through a non-peered VNet. Validation with Threat Hunter and YARA before anything touches production. And a ladder, not a cliff: with the VBR you already own covers every workload today, and VRO turns the workloads that matter most into a rehearsed, documented, one-click routine. Define your minimum viable company. Pick your immutable copy. Build the isolated network, it takes less time than reading the other guys' deployment guide. Then practice until the answer to "can we recover?" is "we already have."

  • Anthropic Just Wrote the Case for Data Resilience (They Just Didn't Call It That)

    Anthropic published a security framework called “Zero Trust for AI Agents” in May. If you build, deploy, or secure AI agents, read the whole thing. It is one of the clearest pieces of guidance on agent security I have seen, and it comes from a leader in the space. I spend my days (and nights) thinking about data resilience and data security, so I read it asking a simple question. Where does this framework actually touch the things my world cares about? I expected a few loose connections. Instead, multiple sections read like I slipped the author $20 to write Veeam's positioning. That is not a coincidence. The direction this framework points is the direction enterprise security has been moving for years, and it is the direction Veeam has been pointing toward for the AI Era. Zero Trust for Agents Is a Different Animal Zero Trust is not new. The phrase goes back decades and the principles were codified by NIST and later the NSA. Never trust and always verify. Grant least privilege. Assume breach. None of that is novel. What is novel is the thing being governed. Traditional Zero Trust assumes the entity behaves predictably. A human authenticated through MFA interacts with systems through known workflows. A microservice makes the API calls its code tells it to make. The behavior is bounded. Agents break that assumption, and not because they are flawed. They break it because they are non-deterministic by design. An agent given a tool might use that tool in a way nobody anticipated. They reach into databases, APIs, and tools through MCP integrations. They chain those actions together at machine speed. The blast radius of one compromised agent is far larger than one compromised user account, and it expands the moment you connect another tool. Anthropic is blunt about the stakes. They write that frontier models are compressing the timeline between vulnerability and exploit from months to hours. Threat actors are also leveraging AI to sharpen tactics they have always used, such as reconnaissance, initial access through phishing and vishing, exfiltration, and credential access. The tactics have not radically changed, but the rate at which they execute now runs at machine speed. How the AI Era Is Evolving Each Zero Trust Principle Assume Breach Anthropic is explicit about it. Design your agent deployments for breach from day one. Do not try to prevent every intrusion. Limit the damage when one happens. Segment by identity. Contain and understand the blast radius of each agent. Make sure compromising one system does not hand an attacker the rest. If you have spent any time in backup and recovery, you have heard a version of this your whole career. Assume the bad thing happens. Architect so you can come back from it. This is not groundbreaking. The takeaway here though is that one of the most credible names in AI is now saying it about autonomous agents, in a security framework, to an audience of CISOs and architects who are deploying these systems faster than they are securing them. Least Privilege The framework makes one conceptual move I think will outlast the rest of it. The distinction between least privilege and least agency. Least privilege is familiar. Give an entity only the access it needs. An agent that reads log files should not have write access to production. Least agency goes further. Give an agent only the autonomy it needs for the task in front of it. If it needs to query a database, hand it a parameterized query interface, not raw SQL. If it needs to change a config, give it a scoped API, not shell access. If you accept that an agent will eventually be compromised or simply reason its way somewhere you did not intend, then an agent with narrow agency is a contained incident and an agent with broad agency is a catastrophe. The access controls can be technically correct and the autonomy can still be the thing that hurts you. Never Trust and Always Verify Threat actors are already using AI to move at machine speed, which is widening the coverage gap. The coverage gap is the percentage of alerts that go uninvestigated, and every uninvestigated alert is risk you cannot see. This means we have to move at machine speed on defense too. It does not mean fire humans and replace them with AI. Human in the loop is critical for decision making. It means close the coverage gap with the same kind of automation the attackers are using. An agent for every alert, to triage, enrich, and correlate, so the percentage of findings that actually get investigated goes up instead of drowning a SOC analyst in a queue. Applying Anthropic's Zero Trust Agent Guidance with Veeam Integrity and Recovery Every meaningful control Anthropic recommends here is at the heart of what Veeam does. They tell you to capture a known-good baseline so you can identify a clean state and restore to it when an agent is compromised. That is the entire premise of immutable backup paired with clean restore point identification. You scan, you verify, you know which point in time is trustworthy, and you recover to it. They tell you to segment by identity so a compromise cannot move laterally. That is isolated, immutable storage with its own network and credential boundary. They tell you to run orchestrated response playbooks with graduated escalation. That is exactly how a well-built recovery runbook works. Every decision about network, compute, priority, and restore point is resolved before an incident, so that during the incident a human authorizes and the automation executes. Recovery is the last resort if everything else fails. The question is not only whether you have backups. It is whether you can answer three things under pressure. Are you protecting the data the agents are being fed and acting on? Is that data immutable, so an attacker cannot quietly alter it? And can you map cleanly back to the point of incident, so you know which restore point is actually clean rather than already poisoned? If you cannot answer those, you do not have a recovery posture. You have a prayer. When the most safety-focused AI company in the world tells you to build for breach and recover to a known-good state, that is the data resilience thesis wearing a different badge. Input Validation and Output Controls An agent's output is only as trustworthy as the data feeding it. Anthropic spends real time on this. Memory poisoning corrupts the context an agent uses to make decisions. Tool poisoning tampers with the responses an agent gets back and trusts as fact. RAG pipelines pull from sources that may be stale, over-permissioned, or deliberately tainted. The agent does not know the difference. It reasons over whatever it is given and acts with confidence either way. Strip the jargon and it comes down to one thing. If you cannot vouch for the data, you cannot vouch for the agent. That is a data integrity problem before it is an AI problem, and it is squarely where data resilience and data security posture live. Most organizations are pointing agents at data they have never classified and cannot prove is intact. That is the quiet risk underneath the loud ones. Veeam's DataAI Command Graph shows where your sensitive data sits, who (or what) can touch it, whether it has been altered, and the ability to restore a known-clean version. Observability and Traceability Anthropic wants immutable, append-only audit logs, streamed to a SIEM, and correlated with other security events. Veeam not only meets those benchmarks, it provides an agent activity log that breaks down every action an agent has taken, the users and groups who have access to the agent, and the files and data systems the agent can reach. This is critical for forensic analysis and anomaly detection, but also for compliance. Full audit trails of who touched what data, why, and who authorized it, plus complete lineage from source to output to satisfy the explainability requirements regulators are starting to demand. Veeam provides this through agent activity monitoring which is a log of every action agents took on data. Anomaly Detection Dwell time, how long a threat sits before you detect it, and coverage, the percentage of findings you actually investigate. Anthropic singles these out as the two metrics with the most leverage when exploit windows collapse to hours. Veeam threat detection sets a clean baseline and flags anomalies early, which attacks the dwell time problem directly, while data security posture tooling highlights overly permissive agent access. Agent Authentication and Privilege Management There are two domains where Veeam does not map, and I am not going to pretend otherwise. Agent authentication, and privilege management. Anthropic wants every agent to carry a unique cryptographic identity. They want short-lived tokens issued by an identity provider, just-in-time privilege escalation, per-action continuous authorization, and attribute-based access control (ABAC). These are real, important controls, and they live at a layer that backup and data security tools do not operate in. That is identity provider and platform territory. Why This Matters Now The conversation about securing AI agents has been dominated, reasonably, by the front half of the problem. How do you stop the agent from being compromised in the first place. Identity, least privilege, least agency, prompt injection, sandboxing. That work is essential and it is where most of the attention has gone. But Anthropic's framework, by leaning so hard on assume breach, is quietly making a second argument. You will not stop every compromise. Agents interpret goals, chain tools, and act at machine speed, and at some point one of them will do something you did not intend, whether through manipulation or its own ambiguity. When that happens the question becomes how fast you can identify a clean state and recover. The agent era makes the blast radius bigger and the speed higher, which means the recovery posture matters more, not less. The organizations that come through the next few years in good shape will not only be the ones with the best agent identity controls. They will be the ones who assumed breach, protected and classified the data their agents depend on, kept an immutable known-good state, and could prove they were back to clean. Everyone is racing to stop agents from being compromised. Far fewer are asking what happens when one is. That second question is the one I would be asking before I deployed a single agent against data that matters.

  • Instant Recovery to Azure: Building a Real Cleanroom

    The same three problems come up in almost every cloud conversation I have: meeting RTOs in a cyber event when cloud is the DR location, a need to reduce cloud spend, and identifying a clean restore confidently. This might come as a surprise but Veeam has the answer to all three with the underlying technology that instant restore to Azure provides, yet it is one of the most underutilized capabilities in the Veeam platform. Instant Recovery to Azure is very similar to instant restore to VMware for folks that are familiar. During recovery the VM is live and accessible in Azure before the full data migration to native Azure Managed Disks completes. What makes this unique is that it works for more than just native Azure VMs. Backup files from multiple workloads support instant restore to Azure including: Hypervisors - VMware, AHV, Hyper-V, Proxmox, and HPE Morpheus Public Cloud - EC2, Azure, and GCP Physical Servers - Windows and Linux A key architectural improvement an everyday Veeam user might notice above are Helper Appliance Templates stored in the Azure Compute Gallery. Previously, every restore required spinning up a helper appliance from scratch. Now, OS-specific templates, both Windows and Linux, are pre-published in the Compute Gallery for your target region. The result is a meaningful reduction in time-to-first-boot. What this solves This is great and all Brad but what business problem(s) are we actually solving here? Executives aren't buying 'instant restore.' They're buying confidence. Confidence that the business keeps running, that the breach doesn't make headlines, that the CFO doesn't have to explain a seven-figure cloud bill that grew because nobody right-sized the DR strategy. Let's walk through each of the 3 use-cases. Cyber and Disaster Recovery to Azure: Veeam Vault as a Cleanroom Ransomware recovery is not about 'do we have a backup?' The questions are: Is the restore point clean? Can we validate before restoring to production? Can we stand up a Minimum Viable Business (MVB) fast enough? Do we know the workloads that bring back our MVB? Not only is Veeam Vault an immutable, offsite, logically air-gapped, encrypted copy of your data, but also is the environment where you identify a clean restore point. It's important to note, tight alignment with your security team is non-negotiable here. They own the forensic timeline and the blast radius analysis. Without that, you are guessing at a clean restore point, and that guess has consequences. Veeam has the below Cyber Resilience capabilities though to assist the security team: Veeam's Breach Impact Analysis builds a graph of data accessed and regulations violated per region Recon for Forensic Triage in Incident Response maps common indicators of compromise to the MITRE ATT&CK framework Veeam Threat Hunter and YARA rules to scan backups in Vault to validate it is free of malware indicators The isolation is architectural, not procedural. The restored VMs boot into a VNet with no routes to production. That's the full cleanroom chain: Vault for immutable storage and clean-point identification, isolated Azure VNet for live forensic validation, and instant restore as the mechanism that makes it fast enough to actually meet business RTOs under a real cyber event. Protecting native Azure VMs If your workloads already live natively in Azure, you may assume you're covered by Azure Backup. You are partially right, and that partial coverage is exactly where the risk and costs hide. Native Azure backups only restore the full VM, which requires organizations to store snapshots for a longer retention to achieve better RTOs. Also worth mentioning that Native Azure Backup has no dedupe and compression and therefore requires 2-3x more storage than Veeam. This workaround is costly to the business. You can calculate the cost for yourself here with Microsoft's backup calculator as finding all the different variables in billing is difficult. With Veeam, those image-level backups target Vault for immutable, WORM-protected, logically air-gapped copies. Even if every snapshot in Azure is deleted, the Vault copy is intact. And critically, you can instant-restore those native Azure VMs back to Azure as native VMs in minutes, with Threat Hunter scanning the backup before it boots. Reducing cloud and day 2 spend The total cost of protection is something most vendors hope you do not calculate. In particular, the cost to recover. Compute costs during backup and recovery add up fast, and they don't show up on a data protection vendor quote. They show up on the cloud bill. Architecturally, several vendors have to run their storage on a compute layer. This is far more expensive than storing data on Blob or Vault type media. In addition, vendors that can't restore natively to Azure VMs force customers into one of two expensive paths: restore through an appliance (adding significant compute cost on top of storage), or restore into Azure VMware Solution (AVS) or Nutanix Cloud Clusters (NC2), which are dedicated bare-metal infrastructure that carries a premium price tag and operational overhead to match. Neither path recovers you to a native Azure VM efficiently. Veeam restores directly from Vault, which uses Azure Blob storage, with minimal compute. The workload boots from the backup data and migrates to native Azure managed disks in the background. No intermediary appliance running up a compute bill. Good-to-knows when doing this at home Deploy Helper Appliance Templates before you need them. If your first instant restore attempt is during a live incident and you have not pre-deployed the Compute Gallery templates for that region, you are adding significant time to your RTO while under maximum pressure. This is a quarterly drill item. The isolated recovery VNet is not optional for cyber recovery. I have seen organizations boot a potentially compromised VM into a production Azure VNet to verify it works first. The result is always the same: it made things worse. Build the isolated network. It takes 15 minutes to create. Use it every time. Azure API throttling kills mass restores. Use existing resource groups, NSGs, and storage accounts. Restore point selection is a security decision, not a backup decision. Your security team drives the pre-compromise window using Recon Scanner and forensic analysis. The data protection team executes. These are two different conversations, and they should happen before the incident, not during it. Test quarterly, not annually. Azure changes. Your workloads change. Execute the full recovery test quarterly. Recovery muscle is built through repetition, not documentation. Conclusion The organizations that recover fastest, whether from natural disasters or cyber events, treat recovery as a first-class engineering discipline, not an afterthought. Veeam's instant recovery to Azure, combined with Vault as a cleanroom, provides a recovery chain that is faster, cheaper, and more secure than anything built on native snapshots or competitor appliance models. Define your Minimum Viable Business. Know which workloads land in Azure first. Then practice until the answer to 'can we recover?' is 'we already have.'

  • AI Resilience: A New Era of Data Resilience

    Data protection used to just be about protecting structured data with the occasional operational restore request, but with that question lurking on everyone's mind, "Can I actually do a mass restore in the event of a natural disaster?" Spoiler alert: the answer was always a definitive maybe. As cyberthreats became the new and far more common disaster the industry evolved with it. Data resilience took center stage, and the ability to prove mass and cyber restore capabilities couldn't be a question anymore. It had to be a proven tested outcome. And now, data has become even more critical in the world of AI. Data is the raw materials to AI. If not properly mined and extracted it can be useless or even dangerous. Enter the third evolution of this industry. AI Resilience! AI transforms how we use, share, and protect data. The question isn’t just “Is my data safe?” It’s “Can I trust my data to power AI, meet compliance, and drive innovation without risk?” According to Gartner, only 20% of organizations report getting value from their AI tools and 38% say the cost of AI outweighs the benefit. This is because AI projects are only as good as the data we put into them. That’s where Securiti.ai , now part of Veeam, steps in. Veeam’s acquisition of Securiti.ai is a strategic leap to unify data resilience, security, privacy, and governance into one intelligent platform. Securiti.ai’s Data Command Center is more than a dashboard, it’s a knowledge graph that maps every data asset, user, policy, and risk across hybrid, multi-cloud, and SaaS environments. Imagine answering any question about your data. Where it lives, who can access it, what risks exist, and how it’s being used with full context. What This Means for Existing Veeam Customers: Ever wonder if you're protecting data that doesn't really need to be backed up? Ever find storage spend is far more than it seems like it should be due to outdated retention policies? Ever wonder what data is most critical to bring back a minimum viable business (MVB). Securiti.ai answers these questions and more through the below key capabilities: Reduce Spend and Risk with Redundant, Obsolete, Trivial (ROT) Analysis:  Automatically discover ROT data across structured, unstructured, SaaS and cloud to remove unnecessary data that increase risk and expense to the business. Breach Impact Analysis:  Better prepare for exfiltration based attacks by classifying sensitive (PII, PHI, PCI, etc.) and critical data to better understand your minimum viable business. Regulatory Compliance:  Maps controls to frameworks like GDPR, HIPAA, PCI DSS, and the EU AI Act, generating audit-ready reports to reduce manual effort and risk of fine or penalty What This Means for Veeam+Securiti.ai Customers: Every organization is trying to increase revenue, reduce costs, or/and minimize risk. Each petabyte of data removed saves up to $500k in storage costs, plus eliminates trivial or obsolete data that exposes risk to the business if breached and exfiltrated. It is critical to remediate data ROT. Feeding data ROT into our AI models and pipelines results in garbage in garbage out as validated by a recent Gartner report highlighting 40% of AI projects will be canceled by 2027. Data is the fuel to AI and ensuring accuracy, precision, and contextualized insights of that data is key to organizations successfully adopting AI. Key capabilities of Veeam+Securiti AI that can't be found anywhere else: Comprehensive Discovery and Remediation:  Automatically find shadow, cloud-native, ROT and dark data assets across hundreds of systems—structured, unstructured, and streaming. AI-Powered Classification:  Uses advanced ML and NLP to classify sensitive data (PII, PHI, PCI, etc.) with high accuracy, reducing false positives and negatives. Risk Prioritization:  Identifies “toxic combinations”—ordinary conditions that, when combined, create catastrophic risk (think open payroll data + global access + AI model ingestion). Orchestrated cross-platform cleanrooms:  Orchestrated recovery workflows from natural or cyber disasters cross-platform in an automated, documented and a proven way. The Bottom Line AI is moving fast. Regulators are watching. Legacy tools can’t keep up. Veeam + Securiti.ai delivers the only platform that unifies data resilience, security, privacy, governance, and AI trust, giving you the confidence to innovate, comply, and recover at scale. What do you think? Is AI resilience the future? Comment below!

  • A Silver Lining to Ransomware in 2025 : Dwell Time is Plummeting

    Are you tired of hearing scary ransomware statistics? You’re not alone. It’s an epidemic among vendors, and the fear-mongering can feel relentless. Joking aside, we all know ransomware is a persistent, damaging threat. Let’s skip the usual data points proving its impact and focus on a rare bright spot: dwell time is dropping dramatically. This shift is transforming how organizations combat ransomware, and with Veeam’s storage integrations, it’s easier than ever to turn this silver lining into a recovery superpower. Dwell Time From Months to Hours Dwell time is the period a threat actor lurks undetected in your environment. Attackers would spend months encrypting workloads before dropping the dreaded ransom note. This left organizations in a bind: backups often contained corrupted data, and even if a clean restore point existed from 30+ days ago, losing a month’s worth of data was catastrophic for business continuity. Fast forward to 2025, and Coveware by Veeam reports dwell times shrinking to days or even hours . Why? More frequent audits, behavioral detection, and the rise of Ransomware-as-a-Service (RaaS) toolkits have changed the game. These toolkits follow predictable patterns, compromising systems, moving laterally with living-off-the-land tools, and encrypting or exfiltrating data. While this makes attacks faster, it also means the ability to recover from more than just backup. Veeam: Recovery From More Than Backup So, why is shorter dwell time a win for ransomware victims? When attackers linger for only days or hours, your recovery options multiply. Most organizations maintain multiple Recovery Point Objectives (RPOs) and Recovery Time Objectives (RTOs), such as 24-hour replicas or 3-day storage snapshots, which are far faster than restoring from traditional backups. Veeam unifies these recovery layers into a single, powerful platform. Paired with a mature incident response plan, Veeam delivers a 2x reduction in downtime and 3x faster RTOs, ensuring your business bounces back before the damage spirals. Cutting Downtime with Veeam Storage Integration When ransomware strikes, restoring a minimum viable business (MVB)  as quickly as possible is the goal. Automation and pre-planned scenarios can shave hours if not days. Recovering from storage snapshots is lightning-fast once data is moving, but without orchestration, manual steps create bottlenecks. Imagine your MVB requires 500 VMware VMs, with downtime costing your organization $1 million per hour . With dwell times of 1-2 days, snapshots become your recovery lifeline. Without Veeam, you face: Hours mapping 500 VMs  to their respective datastores. Mounting 30-50 datastores , each taking precious minutes. Manually powering off original VMs  to avoid IP conflicts, registering new VMs, and powering them on. Veeam eliminates this grunt work with seamless storage integrations for Pure Storage, NetApp, Dell, IBM, HPE, and more . Clients are amazed at how effortlessly Veeam: Locates LUNs on the array. Identifies snapshots of those LUNs. Finds VMs within snapshots. Instantly recovers VMs in parallel, no datastore mapping required . This automation transforms a multi-hour slog into a process that takes 15-30 minutes , minimizing downtime and restoring critical operations before the business feels the full sting. Cyber Recovery Done Right In 2025, cyber recovery (CR) is far more common than traditional disaster recovery. Moving data is only half the battle. Ensuring it’s clean is critical. Veeam’s Threat Hunter and YARA rules integrate forensic analysis and isolation into a cleanroom environment, verifying data integrity before migration to production. This automated approach further reduces downtime and accelerates RTOs, ensuring your recovery isn’t just fast but secure. Conclusion: Turn the Silver Lining Into Gold While shorter dwell times offer a silver lining, ransomware remains a formidable threat. The best defense? Be better prepared. According to Veeam’s Data Resilience Maturity Model , 74% of organizations miss key resilience best practices, and over 30% of CIOs overestimate their preparedness. Veeam’s storage integrations and cyber recovery tools automate critical steps, delivering recovery that goes beyond backup. Don’t just survive ransomware. Thrive through it.

  • Veeam: A VMware “Like” Experience with Red Hat OpenShift Virtualization

    Since Broadcom’s acquisition of VMware, Red Hat OpenShift has emerged as a leading platform  for orchestrating containerized and virtualized workloads. With the integration of Red Hat OpenShift Virtualization, organizations can run virtual machines (VMs) alongside containers, enabling a seamless transition from traditional virtualization to modern, cloud-native architectures. However, as enterprises embrace this hybrid approach, ensuring robust data protection, disaster recovery (DR), and ransomware defense is critical. This is where Veeam Kasten   steps in, providing enterprise-grade data management tailored for OpenShift environments.  Why Veeam for OpenShift?   Veeam Kasten is purpose-built for Kubernetes, offering enterprise-grade backup, disaster recovery (DR), and ransomware protection for OpenShift environments. The 2024 Red Hat State of Kubernetes Security Report  highlights that 46% of organizations faced revenue or customer loss due to Kubernetes security incidents, emphasizing the critical need for robust data protection. Kasten addresses this by providing automated, application-centric solutions that scale across hybrid and multi-cloud setups.  Key Features of Veeam:   Comprehensive Backup and Recovery  - Kasten automatically discovers applications and components within OpenShift, protecting stateful workloads like databases. It supports in-node, in-cluster, and cross-cluster backup scenarios, ensuring data is always recoverable, including file-level restores . Regular backup testing, as outlined in the Veeam Kasten and Red Hat OpenShift Virtualization Reference Architecture, helps organizations meet Recovery Time Objectives (RTOs) and Recovery Point Objectives (RPOs).  Disaster Recovery and Orchestration  - Kasten enables DR by replicating backups off-site to meet regulatory requirements. Its transform engine allows seamless migration of applications and VMs between clusters, supporting upgrades (e.g., OpenShift 3.x to 4.x) with minimal downtime. Cloud-native operations with Kubernetes APIs for direct GitOps/ACM integration; Blueprints for robust pre- and post-backup or restore orchestration.   Application Mobility for Red Hat OpenShift Virtualization  - Veeam Kasten supports true application mobility—not just data migration—across OpenShift environments. This includes the ability to move VM-based applications across namespaces, clusters, regions, and even between on-premises and public   cloud OpenShift deployments. For organizations modernizing or consolidating infrastructure, Kasten enables:  Cross-cluster VM migration  for load balancing, disaster recovery, or infrastructure upgrades  Move-to-cloud  workflows, where VMs running on OpenShift Virtualization can be replicated to OpenShift clusters in AWS, Azure, or GCP  Granular workload portability , preserving application configurations, secrets, and persistent volume claims (PVCs)  Hybrid cloud mobility , enabling developers and operations teams to shift workloads dynamically based on business, regulatory, or cost considerations  Ransomware Protection  - Security is paramount in Kubernetes environments. Kasten provides always-on encryption (in-flight and at-rest) using AES-256-GCM keys and supports FIPS 140-3 compliance for stringent security needs. Immutable object storage (e.g., Amazon S3, Azure Blob) protects backups from ransomware, ensuring data integrity. As well as, Cloud-native security design with self-service via Kubernetes-native RBAC. Integration with Red Hat Advanced Cluster Security (RHACS) enhances monitoring, alerting on unauthorized access, such as changes to Kasten’s DR passphrase.  Seamless Integration with OpenShift  - Kasten integrates tightly with Red Hat OpenShift Platform Plus, leveraging OpenShift Data Foundation (ODF) for storage and ArgoCD for continuous delivery. The Kubernetes Backup Solution Brief notes that Kasten’s deployment via the Red Hat OpenShift Operator ensures standardized, repeatable configurations. OAuth authentication aligns with OpenShift’s identity management, simplifying administration.  Real-World Benefits   Hybrid Cloud Flexibility : Kasten supports on-premises, public cloud, and multi-site deployments, enabling organizations to balance cost and performance.   Simplified Operations : Kasten’s multi-cluster manager provides a single dashboard to manage backups across OpenShift clusters, reducing complexity for DevOps teams.  Enterprise-Grade Security : Policies in RHACS, like monitoring Kasten’s DR secrets, protect critical backup infrastructure from threats.  Getting Started with Veeam Kasten   Deploying Kasten on OpenShift is straightforward, using either the Red Hat OpenShift Operator or Helm charts. This solution brief  emphasizes its ease of use, with a modern UI and Cloud Native API that empower DevOps teams. For air-gapped environments, Kasten supports offline installations, ensuring compliance in secure settings.  To learn more, visit Veeam’s OpenShift Backup Solutions  or explore the detailed reference architecture for implementation guidance.  Conclusion   Veeam Kasten delivers a powerful, Kubernetes-native solution for protecting Red Hat OpenShift workloads. By combining backup, DR, and ransomware protection with seamless OpenShift integration, Kasten empowers organizations to modernize confidently while safeguarding critical data. Whether you’re running containers, VMs, or a hybrid mix, Kasten ensures your OpenShift environment is secure, recoverable, and ready for the future.

  • Bouncing Back from a Ransomware Attack to a Minimum Viable Business

    The landscape of business continuity has undergone a significant transformation with the evolution from traditional Disaster Recovery (DR) to the more complex realm of Cyber Recovery (CR). Where DR once focused on mitigating the impact of natural disasters like floods or power outages, CR now contends with targeted cyber-attacks that pose an ever-present threat. For cybersecurity leaders and C-suite executives, the mission is clear: restore a minimum viable business (MVB)—the core operations that sustain viability—amidst chaos. CR requires critical preliminary steps before recovery can even begin. Unlike DR, where recovery is often the easiest part, CR involves properly detecting the incident, containing the attack to prevent its spread, and eradicating the threat actor to avoid double extortion. This multi-layered response underscores why CR poses a significantly tougher challenge than its DR counterpart. Ransomware isn’t just an IT headache—it’s a boardroom crisis. Minimum Viable Business: What Really Happens During a Cyber Recovery Recovery is a strategic race to revive your MVB—those critical systems (e.g., financials, patient records) that keep the business alive. It’s important to understand what actually happens during a CR. Timeline Detective Work : Forensics experts build a timeline to spot affected servers, files, and backdoors—think of it as mapping a crime scene. Veeam’s Recon tool helps pinpoint the scope of the attack and what was impacted. Pinpoint the Safe Spot : Cybersecurity identifies a pre-compromise window (within days) to validate a clean restore before the threat actor was in the environment. This is passed onto the data protection team to validate a clean restore point. Hand It Over : The data protection team takes over, selecting a restore point before the breach. Teamwork here cuts recovery time, targeting key systems to get your MVB running. Cleanroom Check : Backups are tested in a secure cleanroom with Veeam SureBackup, Threat Hunter and any other in-house cybersecurity forensic tools, verifying they’re free of malware. This step ensures no hidden threats derail your MVB. Mass Restore Kickoff : With dwell time going down to single digit days and in some cases hours, Veeam’s multi-layer system—failover replicas, instant recovery from backups, and instant recovery with storage integration—rolls out the fix at scale. It prioritizes restoring your MVB, like bringing online financial or patient records. Fast Recovery Wins : Veeam slashes Recovery Time Objectives (RTOs) with speed. The image shows how antivirus and backup repos team up, ensuring your MVB is operational while full recovery continues. Real Talk: What Works Practice beats panic, especially when aiming for an MVB. Teams that drill with Veeam’s tools recover fastest. Rushing restores without checks can stretch downtime from hours to days, delaying your MVB. Test backups monthly and train your team to spot early signs—panic clicks can wipe evidence needed to prioritize your MVB. The focus on MVB means starting with essentials—say, a retailer’s point-of-sale system or a hospital’s patient intake—before tackling less urgent data. This staged approach, supported by Veeam’s layered recovery, keeps business flowing. A Vision for Data Resilience The future of CR lies in preemptive readiness. Define your MVB today—identify which systems are non-negotiable—and build a culture of resilience. Veeam’s free trial and cleanroom technology offer a starting point, but true leadership means integrating CR into your risk strategy, engaging stakeholders, and staying ahead of evolving threats.

  • 10 Critical Alarms to Stop Ransomware and Protect Your Business Continuity

    Ransomware attacks aggressively target virtual infrastructure like ESXi and vCenter, exploiting vulnerabilities to encrypt data and disrupt operations. Veeam’s advanced monitoring capabilities empower organizations to detect threats early and respond swiftly, safeguarding critical data before backups are compromised. By configuring precise alarms and integrating with SIEM tools such as CrowdStrike Falcon, Palo Alto XSIAM, Rapid7, Microsoft Sentinel, Splunk, etc, Veeam ensures comprehensive security visibility. Below are the top 10 alarms, each mapped to the appropriate MITRE ATT&CK Tactic to fortify your defenses against evolving cyber threats. EDR Services Disabled or Stopped - MITRE ATT&CK Tactic: Defense Evasion (TA0005) Threat actors often disable endpoint detection and response (EDR) services to evade detection. Veeam’s guest OS services alarm triggers alerts if critical defenses, like antivirus or EDR tools (e.g., CrowdStrike, Defender, etc), are stopped, catching evasion attempts early. This prevents ransomware from spreading undetected. Forwarding alerts to SIEM platforms centralizes incident tracking, ensuring rapid response within existing workflows. vCenter Server Brute Force Alarm - MITRE ATT&CK Tactic: Initial Access (TA0001) 45% of cyber incidents Coveware by Veeam responded to in 2024 were targeted at virtual infrastructure. Veeam’s “Bad vCenter Server Username Logon Attempts” alarm detects suspicious login patterns, flagging credential-stuffing attempts that could grant unauthorized access. Sending alerts to incident response platforms enables real-time correlation, strengthening defenses against initial access. ESXi Host Brute Force Alarm - MITRE ATT&CK Tactic: Initial Access (TA0001) ESXi hosts are vulnerable to misconfigurations or exploits, making them attractive social engineering targets. Veeam’s ESXi monitoring alerts unauthorized access and weak setting configurations. Suspicious Ransomware Activity Alarm - MITRE ATT&CK Tactic: Impact (TA0040) Veeam’s suspicious ransomware activity alarm monitors production workloads before backups, a unique feature. It detects anomalies like unusual file changes or encryption attempts (e.g., workload activity spikes), flagging ransomware in real time. Integrating these alerts with threat intelligence, enables swift isolation of compromised systems. Attempted Backup Deletions - MITRE ATT&CK Tactic: Impact (TA0040) Ransomware often targets backups to block recovery. Veeam’s alarm, based on Event ID 41800, “Backup Deletion Attempt Detected,” triggers when unauthorized deletions are attempted. This protects recovery points, maintaining SLA compliance rates. Suspicious File Activity Alarm - MITRE ATT&CK Tactic: Defense Evasion (TA0005) Ransomware may modify files to prepare for encryption. Veeam’s alarm, tied to Event ID 42402 , “Suspicious File Activity Detected,” flags unusual file modifications in workloads. This catches subtle attack indicators. Forwarding these alerts enhance correlation, enabling containment before ransomware spreads. Unauthorized Access Alarm - MITRE ATT&CK Tactic: Privilege Escalation (TA0004) Unauthorized changes to security policies weaken defenses. Veeam’s alarm, linked to Event ID 42402 , “Four-Eyes Authorization Event Created,” detects attempts to bypass four-eyes authorization, requiring dual approval for critical actions. Alerts forwarded ensure oversight, preventing attackers from disabling safeguards. MFA Attempts Exceeded Alarm - MITRE ATT&CK Tactic: Initial Access (TA0001) Excessive multi-factor authentication (MFA) attempts signal probing for access. Veeam’s alarm, based on Event ID 40206 , “MFA Attempts Exceeded,” triggers when login attempts exceed thresholds. This identifies threats targeting MFA-protected accounts. Alerts forwarded help block unauthorized access promptly. Lateral Movement Alarm- MITRE ATT&CK Tactic: Lateral Movement (TA0008) Attackers use lateral movement to escalate privileges, targeting backup servers to disrupt recovery. Veeam’s event-based rule for Event ID 4625, “An Account Failed Logon Attempt,” monitors suspicious logon activity on backup servers. It is best to set the condition to alert after 3 failed logon attempts to limit false positives. This alert ensures rapid investigation and containment. Immutability Change Attempt Alarm - MITRE ATT&CK Tactic: Impact (TA0040) Attackers may alter immutability settings or infrastructure to undermine recovery. Veeam’s alarm, based on Event ID 28100 , “Configuration Change Detected,” flags unauthorized changes to backup immutability or infrastructure settings. Alerts forwarded ensure monitoring of sabotage attempts. Integration and Actionable Insights These alarms align with MITRE ATT&CK tactics—Initial Access, Defense Evasion, Privilege Escalation, Lateral Movement, and Impact—covering critical attack stages. All integrate seamlessly with syslog or SIEM tools, including CrowdStrike Falcon, Palo Alto XSIAM, Rapid7, Microsoft Sentinel, Splunk, etc. For example, an attempted backup deletion (T1490) can trigger a Splunk investigation, while lateral movement alerts (T1078) prompt CrowdStrike Falcon to isolate servers, minimizing attacker opportunities. Veeam’s Data Platform ensure comprehensive protection, detecting everything from initial brute force attempts to final impact efforts like data encryption. Fine-tuning rules to reduce false positives, such as whitelisting known admin accounts, enhances accuracy. Stay vigilant, and let Veeam be your frontline defense against cyber threats.

  • Veeam's Comprehensive Malware Detection: Before, During, and After a Backup

    In the ongoing battle against ransomware, companies are becoming more resilient. Reports from Coveware by Veeam and Chainalysis show a significant decline in ransomware payments in Q4 of 2024. This success is attributed to several key factors, including improved federal regulations, successful takedowns of major cybercriminal groups, and most importantly, organizations being better prepared and more resilient in responding to and recovering from encryption-based malware attacks. Data protection has evolved significantly in recent years. Discussions have shifted from inline vs. post-process deduplication to security focused topics such as workload immutability, incident response, and malware detection. Backups scanning for malware was and never is meant to replace traditional endpoint or extended detection and response tools. It is meant to provide an additional layer of detection for a defense-in-depth approach for detecting malware. Veeam provides the most practical and comprehensive malware scanning capabilities before, during, and after a backup in the data protection industry. Let's explore each use case. Before Backup: Proactive Threat Assessment Based on the MITRE ATT&CK framework and data from Coveware by Veeam, we know that threat actors target backups once they gain initial access. Veeam is the only vendor that proactively looks for suspicious behavior before a backup is taken. Recon Scanner : Provides proactive alerts on potential threats to your backup server, detecting suspicious events like unfamiliar IPs attempting remote access or compromised accounts trying brute force attacks. It helps identify vulnerabilities and builds a timeline of events to pinpoint clean restore points. Observability, Analytics & AI-powered Insights : Detects anomalies in the production environment before a backup, identifying unusual VM patterns, brute force attacks on ESXi and vCenter, and suspicious SSH activity. Veeam Incident API : Allows third-party security tools to integrate with Veeam, flagging potential malicious restore points and triggering out-of-band backups for actively encrypted workloads. During Backup: Multiple Layers of In-line Detection Veeam diligently detects and mitigates threats in real-time during backups, surpassing alternatives that require post-process scanning and metadata be sent to their cloud just to identify basic bulk changes. IoC Scanner : Validates if harmful tools known to be used by cybercriminals are running on machines and detects newly installed tools, even if they are living off the land tools, that can exfiltrate, encrypt or damage data. Entropy Analysis : Scans data blocks for randomness, looking for encrypted data, onion links, and ransom notes. File Indexing : Uses a signature-based analysis to scan Veeam's database for known malware extensions, ensuring quick flagging of potential threats. Immutable Backups : Ensures backups are recoverable from encrypted cyber attacks with options like Veeam Hardened Repository, Veeam Vault storage, and third-party immutability. After Backup: Ensuring Fast and Clean Recovery Unfortunately, ransomware wouldn't exist if prevention and detection tools caught everything. Organizations must prepare for the worst but hope for the best. Post-process scanning is crucial to ensure clean data restoration and avoid reinfection if a threat actor evades detection. Recon Blast Radius : Identifies the actual scope of a ransomware attack, detecting corrupted or non-encrypted files and building a timeline of events. Threat Hunter : Scans restore points for malware using a signature-based antivirus engine, ensuring only clean data is used for recovery. YARA Rule-based Scanning : Looks for indicators of compromise, detecting malware that might have been missed by other tools or is a zero-day Orchestrated Restore & Cleanroom Capabilities : Creates detailed recovery plans, testing and validating the recovery of critical applications. Putting it All Together These scanning capabilities are most effective when security teams are involved. Forwarding events to the tools and dashboards your security team already uses enables automated processes. Let's put this all together in an example to better visualize how all these components work together: Veeam or an EDR/XDR tool detects suspicious behavior on a machine before or during a backup. Event is forwarded to the organization's SIEM tool. A playbook automatically kicks off that triggers an instant restore of the suspected infected machine to a cleanroom environment for a second opinion from Veeam Threat Hunter. If the scan comes back clean, it marks the event as a false positive. If the scan confirms malware, Recon Blast Radius scans the machines to better understand the timeline and scope of impacted data. Finally, don’t be stuck just recovering from spinning disk backup. Recover from more than just backup once the scope of the attack is known and the threat actor is properly eradicated. Conclusion Veeam's commitment to cybersecurity extends throughout the entire data lifecycle. From before backup, during backup, and after backup. Veeam ensures that your data is protected at every stage. This comprehensive approach not only minimizes the risk and impact of cyber attacks but also ensures that organizations can recover quickly and efficiently. In conclusion, Veeam's sophisticated scanning technologies and proactive strategies make it the most reliable solution for malware detection and recovery. By integrating advanced tools and maintaining a vigilant approach to cybersecurity, Veeam helps organizations stay one step ahead of cyber threats, ensuring that their data remains resilient and their operations uninterrupted.

  • The Veeam Difference: Why Recovery is Your Currency

    In the ever-evolving landscape of cyber threats, the ability to recover swiftly and effectively from an attack is paramount. Veeam stands out in the field of data resilience by offering comprehensive solutions that ensure businesses can bounce back quickly, minimizing downtime and financial loss. Too often, organizations assess the cost of backup simply by the quote vendors provide for software and hardware. You don't buy backup to backup though. You buy backup to recover. Therefore, it's critical the costs of recovery are factored into an organizations total cost of ownership (TCO) when evaluating vendors. Recovery is Your Currency Despite all the investments in detection and prevention tools, these attacks are still prevalent and frequently occurring. Cyber resilience is about more than just preventing attacks; it's about being prepared to recover when they inevitably occur. When you fallback on your last resort (backups), you want to be able to restore confidently, quickly and in a cost-effective manner. Many cyber recovery and disaster recovery costs are not forecasted properly and take businesses and budgets by surprise once it is too late. Several recovery costs that organizations need to factor into their overall TCO when selecting a backup vendor are: An additional DR tool for critical workloads instead of a solution that includes both backup and DR to minimize downtime Running VMware on cloud because the backup tool cannot convert a hypervisor or physical machine backup into a native cloud workload Most importantly the compute cost to constantly run an appliance in the cloud for offsite data instead of replicating to native object storage See below for a more detailed breakdown of a real-life enterprise scenario: The Veeam Difference: Two Responses to an Attack In the event of a cyber attack, costs are the last thing on your mind though. The speed and efficiency of your recovery process can make all the difference. Veeam's approach ensures minimal material damage by identifying the scope of the attack and minimizing downtime with fast orchestrated recovery options. Here's a comparison of responses to an attack with and without Veeam: Without Veeam : Companies may face significant delays in detection and recovery, leading to greater damage and higher costs. Threat actors can move laterally and gain access to critical systems without notification, resulting in prolonged downtime and limited restore options. With Veeam : Veeam detects initial access, lateral movement to the backup server and encryption/exfiltration activity inline during the backup, sending critical alerts to the security team. This enables quick diagnosis, minimal data encryption, and guided negotiations to end downtime swiftly with included incident response services with Coveware by Veeam. The result is minimal material damage and fast recovery. Conclusion In today's digital age, recovery is indeed your currency. Veeam's comprehensive approach to cyber resilience, advanced recovery options, and proactive threat management make it the ideal choice for businesses looking to protect their data and ensure business continuity. With Veeam, you can confidently navigate the complexities of cyber threats and emerge stronger.

  • Scan Cloud Workloads for Malware with Veeam

    Ransomware attacks continue to evolve, targeting cloud-based workloads and posing significant risks to organizations. These malicious attacks can encrypt critical data, disrupt operations, and demand hefty ransoms for data decryption. Not to mention double extortion and the risk of stealing/selling the victims data. To counter such threats, organizations must prepare robust defense strategies to safeguard their cloud workloads. Part of that defense strategy not only is backing up those cloud workloads but also choosing a solution that follows a defense-in-depth strategy. Veeam's inline scanning capabilities provide organizations with real-time threat detection, malware protection, compliance adherence, and data loss prevention. By incorporating Veeam into your cloud infrastructure, you can establish a backup security posture and ensure the safety and integrity of your cloud-based workloads. As a result, businesses can operate with confidence, knowing that their critical data and operations are protected against evolving cyber threats. Veeam has two types of inline scanning: Inline entropy analysis at the block level File system activity analysis indexing at the guest level Inline entropy analysis looks for files encrypted by malware and anomalies in the data.  For example, if a server in your environment usually has 10% of the data encrypted and changes 50-60 GB a day but all of a sudden we detect that 20-30% of data is encrypted and it is changing 90-100 GB per day that will trigger an alert of suspicious activity on that machine. In addition, you will be alerted if we see machines with ransomware notes, bitcoin addresses or other suspicious content. File system activity analysis is looking for known suspicious extensions from the XML file on the Veeam server. There are over 4,000 known IoCs (indicators of compromise). In addition, this has the intelligence to look for day zero attacks if we start backing up file extensions that have never been seen on a machine even if they are not in the XML file. In the rest of the blog, I'll quickly show how-to setup inline scanning for your cloud workloads. Quick How-To: First, setup a protection group for cloud machines. Simply enter in the credentials to your AWS or Azure account/subscription and choose the region of the machines to protect. Make sure you choose cloud machines for your type of protection group as these are special kind of agents. Select machines individually for the masochists out there or simply insert a tag to catch-all the cloud workloads to protect. A role with the following permissions needs to be attached to each instance you want to protect in the case of AWS. Veeam automagically can create and assign those roles which makes life a lot easier for organizations with hundreds or thousands of workloads out there. Lastly, enable both entropy and file system analysis in the global settings as shown above earlier. Once your backup job is created, inline malware scanning will occur by default on all backups. If something suspicious is detected it will show in the properties of the job, inventory menu and/or syslog server if you enabled event forwarding. And here it is in the properties of the job if you want to get a quick glance of the last restore point Veeam believes to be clean. Conclusion: As the ransomware threat landscape continues to evolve, organizations must remain vigilant to protect their cloud workloads. Veeam's ransomware inline scanning capabilities offer real-time threat detection, malware signature recognition, behavior analysis, and automatic remediation. By incorporating Veeam into your data protection strategy, you can improve the resilience of your cloud workloads against ransomware attacks, ensure business continuity, and safeguard your critical data assets. Don't wait until it's too late – take proactive measures to defend your cloud environment today. Note: Cloud workloads are supported for AWS and Azure Windows IaaS workloads. Entropy analysis works for volume-level backup mode. File system analysis can be used for both image and volume-level backup modes.

  • Forward Critical Veeam Events to Splunk

    Monitoring and analyzing events from various sources is crucial for maintaining a secure and efficient IT environment. Veeam v12.1 offers comprehensive event logging capabilities for specific security related events to help monitor for potential cyber threats. By forwarding Veeam events to Splunk, organizations can take advantage of Splunk's robust log management and analysis features for enhanced visibility and troubleshooting. Below are a few key benefits of forwarding Veeam events to Splunk or any syslog server. Centralized Log Management: By consolidating Veeam events into Splunk syslog, organizations can have a centralized repository for all their logs, making it easier to search, analyze, and correlate data from multiple sources. Real-Time Monitoring: Splunk provides real-time alerting capabilities based on specific events or patterns. By forwarding Veeam events to Splunk, organizations can set up proactive alerts for early detection and remediation. Advanced Log Analysis: Splunk's powerful search and analysis features enable organizations to gain deep insights into their Veeam events. With the ability to create dashboards, reports, and visualizations, you can easily identify trends, patterns, and potential issues within your backup infrastructure. How to setup on Splunk: Launch the Splunk web interface and navigate to "Settings -> Data Inputs". Select "UDP" or "TCP" under the "Syslog" category, depending on your preferred protocol. Configure the port number 514. Can use 6514 over TLS if preferred Only accept connections from VBR How to setup on Veeam: Simply go to "Options" under global settings (hamburger helper in top left) and "Event Forwarding." Enter in the Splunk server and protocol/port to communicate with. Create Splunk Alerts for Veeam Critical Events: By default Veeam forwards all events to the syslog server which can quickly become overwhelming. Luckily, it's easy to create alerts for specific critical events that would matter most to the security team. Below are just a few examples of Veeam events that are important to forward: 42402 - Four-eyes authorization request initiated 42402 - Attempted deleted backup 42220 - Restore point marked as infected 41600 - Malware activity detected 40205 - Invalid MFA code 150 - Time shift detected on repository These alerts can easily be setup in Splunk by searching for the above instance IDs and then creating alerts for them. Below are several alerts I created in Splunk to help filter out the noise for specific events security teams should be aware of. This functionality would work on any Syslog server. By forwarding Veeam events to Splunk, organizations can achieve centralized log management, real-time monitoring, and advanced log analysis capabilities. Splunk's powerful search and analysis features can help identify potential issues, track trends, and optimize backup infrastructure performance. With this integration in place, IT teams can proactively address concerns, ensure data protection, and minimize the impact of backup-related incidents on overall business operations.

Subscribe Form

bottom of page