Most Manufacturing Downtime Has Nothing to Do with Failed Disks, So Why Are We Still Relying on RAID? | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #ransomware


RAID, or the “Redundant Array of Independent Disks” earned its place in enterprise and industrial computing for a good reason. In the 1980s and 1990s, when hard drive failures were a far more common operational headache, RAID solved a genuine problem for companies. The idea was that by mirroring or distributing data across multiple disks, systems could continue running even when individual drives failed unexpectedly – which happened more than you might think. It was an elegant answer to one of the most disruptive events an organization could face at the time. For manufacturers, OEMs, and industrial operators whose systems needed to remain available around the clock, RAID became synonymous with reliability.

But RAID’s success helped to cement a misconception that persists to this day. At its heart, it’s a hardware availability mechanism, but many organizations started thinking of it as part of their dedicated backup strategy, particularly in embedded systems, industrial automation platforms, diagnostic equipment, kiosks, edge devices, and OEM products deployed into customer environments. The problem is that the risks facing these systems in the 1990s and the risks facing those systems today are very different. Today, many of the incidents that disrupt industrial systems aren’t caused by failed disks, but by corrupted updates, ransomware, software failures, accidental deletions, configuration errors, and countless other issues that RAID was never designed to address.

That begs the question, if the failures that actually disrupt operations have evolved, why are so many resilience strategies – particularly in manufacturing – still so heavily reliant on a feature that can’t guard against those failures? Customers certainly aren’t asking whether a disk survived. They care whether production stopped, whether service engineers need to be dispatched, or how quickly a machine can be returned to operation when something goes wrong.

Ransomware doesn’t care how many disks you have

Here’s a thought experiment for manufacturers. Imagine a critical machine suddenly goes offline and production stops. Operators can’t access the system, and engineers are called out while customers wait patiently for updates. What are the chances the root cause was a failed hard drive? In most cases, virtually nil. Disk failures remain possible, but they’re no longer the primary cause of many operational outages, particularly compared to the spinning disks that dominated industrial computing a generation ago.  The incidents that cause the most disruption today are far messier than simple hardware failures, often affecting the entire operational environment rather than a single component.

RAID wasn’t designed to protect against these types of failures. In fact, it often does the exact opposite, because it was built to preserve, and in doing so, it preserves the problem. If ransomware encrypts a file, RAID faithfully mirrors the encrypted version. Similarly, if a software update corrupts the operating system, every mirrored disk contains the same corrupted operating system. This is why there needs to be a clear distinction between availability and recoverability. RAID takes care of the former, but does almost nothing to help with the latter – and in some cases, makes things worse. According to reporting from SC World, only 29% of organizations currently employ layered ransomware protection for their backups, highlighting just how many environments remain vulnerable even before RAID-only strategies enter the picture. The truth is that many systems have been engineered to survive the least likely failure while remaining exposed to the incidents most likely to occur.

The machine is no longer the data center

One reason the RAID conversation has persisted for so long is that many resilience strategies were originally developed for environments where systems remained close at hand. If something went wrong, an IT team could investigate the issue, access the hardware, and begin recovery almost immediately. But that’s rarely the reality for today’s OEMs. The moment a machine leaves the factory floor and arrives at a customer site, control begins to disappear. Industrial equipment may spend years operating in a manufacturing plant, warehouse, laboratory, utility facility, or transportation hub with little direct oversight from the people who designed it. And in many cases, the only indication that something has gone wrong is a support ticket, a frustrated customer, or a production line that has suddenly ground to a halt. That’s not a great position to be in. 

In many ways, the economics of failure have changed. A corrupted operating system or damaged application is no longer a technical inconvenience that can be resolved with a quick visit to the server room. It can trigger engineer callouts, lengthy troubleshooting exercises, replacement hardware shipments, and costly downtime while the root cause is identified. For OEMs, the financial impact often extends well beyond the immediate service costs. Every hour spent recovering a customer system is an hour in which confidence in the product is being stress tested. Customers don’t tend to judge reliability by the elegance of the underlying architecture or the sophistication of the redundancy built into a device. When a machine fails, they simply want to know how quickly it can be returned to operation. As industrial systems become increasingly software-driven, that’s forcing manufacturers to rethink the idea of reliability – is it about preventing failures? Or is it more about how effectively they can recover from them? Ideally, it’s both. 

From mirroring disks to restoring systems

Some manufacturers and OEMs are indeed beginning to change their approach to reliability. Instead of asking how to keep a system running through a hardware failure, they’re asking how quickly an entire machine can be returned to a known-good state after any failure. Image-based recovery has emerged as one potential answer because it captures a complete point-in-time image of the operating system, applications, drivers, configurations, and settings, allowing organizations to simply restore an entire machine to a known good state rather than painstakingly rebuilding it piece by piece. Traditional resilience principles such as the 3-2-1-1-0 methodology, which advocates maintaining multiple copies of data across different media with offline or immutable protection and zero backup errors, still have an important role to play, but manufacturers are increasingly discovering that backup strategy alone doesn’t determine recovery outcomes. The location of recovery data, the architecture supporting restoration, and the ability to recover quickly under operational pressure are becoming just as important as how many copies exist. And in an industry where downtime is measured in lost production, delayed shipments, and damaged customer relationships, the ability to restore a machine predictably and consistently is beginning to matter more than the redundancy of the storage sitting inside it.

Manufacturing has always been an exercise in managing uncertainty. Components will always fail, software will behave unexpectedly, people will make mistakes, and cybercriminals will continue to find new ways into operational environments. RAID remains a valuable technology for addressing one of those risks, but it was never intended to shoulder the weight of modern resilience on its own. Modern resilience isn’t just about keeping systems available during a hardware failure, but ensuring they can be recovered quickly from the incidents that actually cause downtime. What matters now is having confidence that when disruption arrives systems can be restored quickly, consistently, and without unnecessary complexity. After all, when production stops, nobody asks how many disks were mirrored. The only question that matters is how long it will take to get back up and running.

 

Join our LinkedIn group Information Security Community!

——————————————————-


Click Here For The Original Source.

National Cyber Security

FREE
VIEW