The vulnerability traces back to a Linux device mapper configuration that prioritized rapid provisioning over the complete erasure of blocks between customer sessions. In the high-stakes environment of edge computing, where milliseconds determine the success of serverless functions and containerized workloads, infrastructure providers often seek optimizations that minimize latency during cold starts. This specific optimization, however, inadvertently created a side channel for data leakage between different customers sharing the same physical hardware. The discovery highlights a recurring challenge in cloud security: the inherent tension between the speed required for modern ephemeral workloads and the absolute isolation required for multi-tenant safety. As organizations increasingly migrate sensitive database operations and AI reasoning tasks to the edge, the integrity of the underlying storage stack becomes a critical point of failure that bypasses traditional network-layer defenses and identity management protocols.
The implications of this flaw extend beyond a simple software bug, touching upon the foundational trust models that underpin the entire serverless ecosystem. Customers typically assume that once a workload terminates, its associated memory and storage footprints are wiped clean before the next user occupies that space. This incident proves that such assumptions can be undermined by low-level system configurations that favor throughput. While Cloudflare has moved quickly to address the issue, the disclosure serves as a wake-up call for security teams to scrutinize the “magic” of cloud automation. The complexity of modern Linux kernels and the various layers of abstraction used to manage storage at scale mean that even minor configuration changes can have catastrophic security outcomes. This situation necessitates a deeper look into how data is actually handled when it is discarded in a multi-tenant environment.
1. Public Disclosure: The Official Announcement
The security community first learned of the specific risks involving the Cloudflare Containers platform on September 24, 2026. This announcement followed several weeks of internal investigation and remediation efforts triggered by a report from an external security researcher. According to the company’s official communication, a flaw in the underlying storage management system allowed one paying customer to potentially access data remnants left behind by another customer’s previous workload. This disclosure was particularly striking because the vulnerability did not rely on traditional exploit vectors like phishing, compromised credentials, or software exploits. Instead, it was a structural weakness in how the platform allocated and reclaimed shared physical disk space, making it a “passive” but highly effective means of information gathering for an opportunistic attacker.
Despite the severity of the potential data exposure, the bug did not receive a standard CVE identifier at the time of the public writeup. This absence of a tracking number is notable in the current cybersecurity landscape, as enterprise security teams often rely on these identifiers to automate their risk assessments and patching schedules. The lack of a CVE suggests that the issue was viewed primarily as a provider-side configuration error rather than a vulnerability in a specific, distributable software package. Nevertheless, the technical details provided in the disclosure were comprehensive, illustrating exactly how a specific write operation could be used to harvest uninitialized storage blocks. This level of transparency was intended to reassure the market, though it also provided a stark look at the vulnerabilities lurking within shared cloud infrastructure.
2. Sequence of Events: Timeline From Detection to Resolution
The journey from the initial discovery to the final global cleanup spanned approximately three weeks in September 2026. It began on September 4, when Oren Yomtov, a researcher at the security firm Accomplish, submitted a report through Cloudflare’s bug bounty program on the HackerOne platform. This initial report provided the necessary proof-of-concept to demonstrate how sensitive data could be recovered from uninitialized blocks. Recognizing the gravity of the situation, the engineering teams initiated an immediate review of the container provisioning logic. By September 7, just three days after the initial report, a primary mitigation was successfully deployed across the fleet to block the known exploit path, effectively stopping the immediate threat of active exploitation by new workloads.
Following the initial fix, the focus shifted to a much more labor-intensive phase: the total remediation of the existing infrastructure. While the configuration change prevented future leakage, it did not automatically clear the data remnants already present on shared disks. Between September 7 and September 19, the company undertook a global effort to retire and rebuild affected container disks and cached image snapshots. This process was completed on September 19, ensuring that all storage blocks across four continents were either zeroed out or replaced with fresh allocations. The final public statement and technical post-mortem were released on September 24, after the engineering teams had verified that the original proof-of-concept was no longer functional and that the environment was once again secure for multi-tenant use.
3. Mechanism: Multi-Tenant Storage Isolation Standards
To achieve the massive scale required for modern cloud services, providers rely on technologies that allow thousands of different customers to share the same physical servers and storage arrays. The most common tool for managing this in Linux environments is the device mapper thin provisioning system, often referred to as dm-thin. This system allows for “over-provisioning,” where the platform can promise more storage to users than is physically available, only allocating actual blocks of data when a user writes to them. This efficiency is what allows serverless containers to spin up in milliseconds, as the system does not need to pre-allocate massive, empty files for every new session that starts and ends.
In a secure multi-tenant environment, the transition of a storage block from one user to another must be strictly controlled. When a customer finishes a task and the storage block is returned to the “free” pool, the system must ensure that the next user cannot see what was there before. Standard security protocols require that these blocks be “zeroed out”—meaning every bit of data is overwritten with zeros—before they are reassigned. This erasure process is the primary barrier preventing cross-tenant data leakage. If this step is skipped or optimized away to save on disk I/O cycles, the isolation between customers effectively vanishes at the hardware level, regardless of how strong the network-level encryption or application-level permissions might be.
4. Technical Analysis: The skip_block_zeroing Configuration
The technical root of the Cloudflare vulnerability was traced to a specific configuration flag within the dm-thin system known as skip_block_zeroing. This setting is often used in performance-sensitive environments because it eliminates the overhead of overwriting storage blocks before they are assigned to a new workload. By bypassing this step, the system can provide a “new” block to a container almost instantly. However, in a multi-tenant cloud setting, this is a dangerous shortcut. Cloudflare’s implementation used a 64 KB block size for its thin-provisioned pools. When a new container requested storage, it was handed one of these 64 KB blocks which, due to the configuration, still contained the data from the previous tenant who had used that physical space on the disk.
The exploit became possible when a container performed a partial write to a newly assigned block. Specifically, if a user wrote exactly 4 KB of data to the start of a fresh 64 KB block, the system would mark the entire 64 KB region as “allocated” to that user. Because only the first 4 KB were actually overwritten by the new user’s data, the remaining 60 KB remained untouched from its previous state. The new tenant could then perform a raw read of the entire 64 KB block, effectively recovering 60 KB of a stranger’s data for every 4 KB they wrote. This mathematical imbalance—where a small, legitimate action triggers a large, illegitimate exposure—is what made the flaw so potent and easy to trigger for anyone with the technical knowledge of the underlying storage architecture.
5. Consequences: Impact of Partial Writes on Data Exposure
The data recovered through this storage flaw was not merely random noise or encrypted gibberish; it consisted of high-value, structured information. During the internal audit conducted by the engineering teams, they discovered that these 60 KB remnants often contained directory structures, file system metadata, and sensitive configuration files. Most alarmingly, the researchers found that in some instances, they could recover fully intact SQLite database files or significant portions of database pages. This type of information is a goldmine for attackers, as it can contain session tokens, user lists, hashed passwords, or internal application logic that can be used to launch further, more targeted attacks against the original owner of the data.
Because the leaked data came directly from the physical disk blocks, it bypassed all of the standard security controls that organizations put in place. Encryption at rest often fails to protect against this specific type of leakage if the cloud provider handles the encryption keys at the block level and decrypts the data as it is read by the “authorized” (though in this case, incorrect) tenant. The audit revealed that 18 out of 24 tested placements showed some form of residual data from foreign workloads. This high frequency of exposure suggests that the problem was not a rare occurrence but a systemic byproduct of how the platform operated under heavy load, where storage blocks were constantly being recycled between different customers’ ephemeral containers and worker functions.
6. Shared Infrastructure: Vulnerability Spread and AI Risks
One of the more concerning aspects of this disclosure was that the vulnerability was not limited to the standard Cloudflare Containers service. Because Cloudflare Sandboxes—a product designed for running untrusted code, such as AI-driven agents or code interpreters—was built on the same underlying infrastructure, it inherited the identical storage flaw. This meant that any organization using the Sandbox environment to process sensitive data or run automated AI reasoning tasks was just as vulnerable as those running standard containerized applications. The shared nature of the storage pool across different product lines created a wider blast radius than many customers likely realized when they originally signed up for these services.
The rise of AI agents has made this type of infrastructure vulnerability particularly relevant. AI agents often operate by spinning up temporary sandboxed environments to execute code or process data snippets. If these sandboxed environments are susceptible to data remnant leakage, an AI agent could inadvertently “learn” or expose sensitive data from a completely different customer’s session simply by performing routine storage operations. This adds a new layer of risk to the deployment of AI at the edge, where the isolation of the execution environment is the primary defense against data poisoning and intellectual property theft. The discovery that even specialized “sandbox” products were affected underscores the difficulty of maintaining true multi-tenant separation across a diverse and rapidly evolving product ecosystem.
7. Detailed Findings: Scope of the Internal Audit
To understand the full extent of the risk, an exhaustive internal audit was conducted across the global network. The investigation involved sampling production placements and inspecting every reachable directory block to see if foreign data was present. The findings were sobering: affected nodes were found on four different continents, indicating that the skip_block_zeroing configuration was a standard part of the global deployment template. Out of the total directory blocks examined—numbering over five thousand—the audit identified 2,700 distinct foreign directory inodes. This meant that nearly half of the examined blocks contained evidence of data belonging to a tenant other than the one currently assigned to the block.
The audit also highlighted the geographic reach of the problem. It was not a localized issue affecting a single data center or region but was instead prevalent across the majority of the tested infrastructure. Specifically, 20 out of 22 underlying servers checked showed evidence of cross-tenant residue. These statistics confirmed that the vulnerability was a pervasive feature of the platform’s storage architecture. For enterprise customers, these numbers are particularly concerning because they suggest that the likelihood of their data landing on a “dirty” block—or their data being left behind on one—was statistically high during the period the configuration was active. The sheer volume of foreign inodes recovered points to a massive, albeit unintentional, data mixing event that lasted until the remediation was finalized.
8. Barriers: Practical Exploitation and Target Limitations
While the vulnerability was technically simple to trigger, there were significant practical barriers that prevented it from being easily used for targeted corporate espionage. To exploit the flaw, an attacker first needed a Workers Paid account and the ability to deploy complex container workloads. More importantly, the way Cloudflare manages workload placement acted as a natural, though unintentional, defense. Workloads are distributed automatically across a massive global fleet based on capacity and latency requirements. An attacker had no way to choose which physical server their container would land on, nor could they predict whose data might have occupied that server’s storage blocks previously.
This randomness meant that any attempt to harvest data was essentially a “fishing expedition” with no guarantee of whose information would be caught. An attacker could not say, “I want to steal data from Company X,” because they could not force their container to follow Company X’s container onto the same physical disk. This limitation significantly reduced the utility of the bug for state-sponsored actors or competitors looking for specific intellectual property. However, it did little to protect against opportunistic attackers who were simply looking for any valid credentials, database fragments, or private keys they could find. For these actors, the ability to harvest random data from a high-volume cloud provider is still a highly valuable capability, even without a specific target in mind.
9. Remediation Strategy: Global Cleanup Efforts
The remediation of the storage flaw was a two-tiered process that required both immediate software changes and a long-term infrastructure overhaul. The first step was the straightforward disabling of the skip_block_zeroing flag. This change ensured that all future storage allocations would be properly zeroed out before being handed to a container, effectively closing the window for any new data leakage. However, this did not solve the problem of the “dirty” blocks already in circulation. To address this, the engineering teams had to systematically identify every active storage volume and image snapshot that had been created while the flawed configuration was in place and remove them from the system.
This global cleanup was a massive logistical undertaking. It required retiring disks and deleting cached snapshots across the entire global network without causing service interruptions for customers. The process took twelve days of continuous work to ensure that every single block of potentially contaminated storage was either overwritten or physically decommissioned. Cloudflare confirmed that customers did not need to take any action themselves, as the fix was applied entirely at the infrastructure layer. This “invisible” remediation is standard for provider-side flaws, but it also means that customers are entirely dependent on the provider’s internal auditing and execution to ensure that the cleanup was truly comprehensive and that no pockets of residual data were missed in remote data centers.
10. Persistence: Uncertainty Regarding Historical Exploits
A critical point of concern in any data breach or vulnerability disclosure is whether the flaw was exploited before it was discovered. Cloudflare has stated that their review of retained telemetry and logs showed no evidence of unauthorized access or exploitation by third parties. They only identified the testing performed by the reporting researcher and their own internal verification runs. While this is a positive sign, it comes with the reality that standard platform logging is often not designed to detect the specific types of raw disk reads used in this exploit. If an attacker had been quietly harvesting data remnants for months, their activities might have appeared as normal, high-I/O container workloads in the existing telemetry.
Because of these logging limitations, the company cannot offer a one-hundred-percent guarantee that no data was ever compromised. This uncertainty is a common theme in hardware-level or low-level infrastructure bugs, where the “crime scene” is the raw storage medium itself rather than an application-level log file. For organizations that processed highly sensitive information during the period the vulnerability was active, the lack of “evidence of exploitation” must be weighed against the technical possibility that their data was readable by others. This situation highlights the importance of end-to-end encryption, where data is encrypted before it ever reaches the cloud provider’s storage blocks, providing a final layer of defense even when isolation at the infrastructure level fails.
11. Comparison: Data Remnants Versus Container Escapes
To understand the unique nature of this flaw, it is helpful to compare it to other high-profile cloud security incidents, such as “container escapes.” In a container escape, like the Azurescape incident or the GKE Fragnesia flaw, an attacker finds a way to break through the digital walls of their container to gain access to the host operating system or other tenants’ environments. These are “active” attacks that involve bypassing complex security boundaries. In contrast, the Cloudflare storage flaw was a “passive” failure of cleanliness. The container stayed perfectly within its assigned boundaries; it simply found that the “room” it was given had not been cleaned after the previous occupant moved out.
This distinction is important for risk modeling. While a container escape often gives an attacker broad, active control over a server, a data remnant flaw is more about information disclosure. An attacker using the storage flaw does not necessarily gain the ability to execute code in another customer’s environment, but they can still steal the secrets needed to log into that environment later. The Cloudflare incident demonstrates that even if a container runtime is perfectly secure and unhackable, the shared resources it relies on—like storage and memory—can still provide a path for cross-tenant interference. The focus of cloud security is shifting from just preventing escapes to ensuring that the entire lifecycle of a tenant’s resources, from allocation to destruction, is handled with equal rigor.
12. Market Impact: Competitive and Trust Consequences
The disclosure of a cross-tenant vulnerability on a major platform like Cloudflare has inevitable ripple effects in the cloud marketplace. Cloudflare has been aggressively positioning its Containers and Workers products as a faster, more modern alternative to established services like AWS Fargate or Google Cloud Run. This incident gives competitors a significant talking point, particularly when pitching to risk-averse enterprise customers in the finance, healthcare, and government sectors. These industries have strict compliance requirements, such as SOC2 and HIPAA, which mandate rigorous data isolation. A public admission that storage blocks were not being zeroed out between customers could lead to a temporary slowdown in adoption among these high-value clients.
However, the impact may be mitigated by the way Cloudflare chose to handle the disclosure. By providing a detailed technical post-mortem and being transparent about the scope of the audit—including the number of affected nodes and continents—the company demonstrated a level of maturity that is sometimes lacking in the industry. Many cloud providers have historically been opaque about their internal infrastructure failures. Cloudflare’s willingness to “show the math” may actually build long-term trust with the technical community, even as it creates short-term challenges for the sales team. The incident serves as a reminder that in the cloud era, security is not a static feature but a continuous process of discovery, remediation, and transparent communication.
13. Security Recommendations: Guidance for Compliance Departments
For security and compliance teams, this incident provides several clear lessons for managing vendor risk in a multi-tenant environment. First, organizations should conduct a retrospective risk assessment to determine if any highly sensitive data was processed in Cloudflare Containers or Sandboxes during the risk window, which ended on September 19, 2026. If sensitive keys or unhashed user data were stored on local container disks during that time, a rotation of those secrets may be a prudent precautionary measure. This proactive approach is better than relying solely on the provider’s statement that no evidence of abuse was found, given the inherent gaps in telemetry for this type of low-level access.
Looking forward, procurement and security teams should update their vendor questionnaires to include specific technical questions about storage lifecycle management. Instead of asking generic questions about “data isolation,” teams should ask how the provider ensures that block-level storage is zeroed out or cryptographically erased between tenant reassignments. Furthermore, this incident reinforces the best practice of avoiding the storage of long-lived secrets or sensitive databases on ephemeral container disks. Using external, dedicated secret management services and ensuring that all data written to local disk is encrypted with customer-managed keys can render the contents of a “dirty” storage block useless to any future tenant who might happen to find it.
14. Future Trajectory: Evolving Standards for Cloud Isolation
The fallout from the Cloudflare storage vulnerability indicated a shift in how the industry approached the security of thin-provisioned resources. In the months following the disclosure, other major cloud providers quietly initiated their own audits of dm-thin and copy-on-write storage configurations to ensure they were not making similar trade-offs for the sake of performance. This led to a broader conversation among infrastructure engineers about the necessity of hardware-accelerated zeroing, which allowed for the rapid erasure of data blocks without the heavy CPU overhead traditionally associated with the process. The incident proved that as workloads became more ephemeral, the speed of “cleaning up” became just as important as the speed of “starting up.”
Regulatory bodies and industry groups also took notice, with some proposing updates to security frameworks that specifically addressed the recycling of physical resources in shared environments. The bug bounty market responded as well, with increased payouts for researchers who could demonstrate novel isolation failures at the storage or memory layers. Ultimately, the industry moved toward a “zero-trust” model for internal infrastructure, where the assumption that the underlying host was clean was replaced by automated verification and more aggressive erasure policies. While the Cloudflare flaw was a significant setback at the time, it catalyzed a series of improvements that made the entire serverless ecosystem more resilient against the risks of shared infrastructure. Past events showed that the only way to maintain trust in the cloud was to prove, through rigorous auditing and transparent disclosure, that the walls between customers were as solid as the hardware allowed.


