
We Built Redundant Networks. Now We Need Recoverable Networks.
For most of my career, network design has been focused on one primary objective: keep the network up.
We added redundant power supplies, dual WAN connections, backup circuits, high-availability firewalls, redundant switches, secondary data centers, and increasingly sophisticated failover technologies. All of that has been important, and it has made networks much more reliable than they were two decades ago.
But I think there is another question we need to start asking more often. What happens when all of that redundancy still doesn’t solve the problem?
A circuit can fail. A configuration change can take down connectivity. A firewall can lock up. A software upgrade can go sideways. Someone can make a mistake. Sometimes the management network itself becomes unavailable.
When that happens, the question is no longer whether the network was designed with enough redundancy. The question becomes much simpler: how do we get back in and fix it?
That is where I think the conversation around network resilience needs to evolve.
Redundancy Is Not the Same as Recoverability
We tend to use terms like redundancy, high availability, resilience, and business continuity somewhat interchangeably, but they are really different parts of the same problem.
Redundancy gives us alternate components or paths when something fails. High availability is designed to keep applications and services operating through those failures. Network resilience goes a step further and asks how well the organization can continue operating and recover when something unexpected happens.
Recoverability is a big part of that equation. You can have a network with multiple redundant components and still have a very difficult time recovering it remotely.
Take a branch office as a simple example. It may have two Internet connections, SD-WAN, a UPS, redundant switches, and a very well-designed firewall architecture. On paper, that looks resilient.
But what happens if a configuration mistake takes down both WAN paths? What happens if the firewall becomes unresponsive? What happens if an upgrade breaks connectivity?
If the answer is that someone has to get in a car, get on an airplane, call a local employee, or dispatch a third-party technician to reboot something, then there is still a gap in the resilience architecture.
The network may be redundant, but it is not necessarily recoverable.
Remote Access Was the Beginning
I have spent more than 25 years around networking and out-of-band management, and one thing has never really changed. When something goes wrong, engineers need a reliable way to get to the equipment.
Historically, that was the role of the console server. If the production network went down, an engineer could access the console port on a router, switch, or firewall and troubleshoot the problem remotely.
That capability is still very important. But I also think our definition of out-of-band management has become too narrow.
Modern network operations require more than getting a command prompt on a device.
An engineer may need an independent cellular connection because the primary WAN is unavailable. They may need to reboot a locked-up device. They may need to see whether a circuit is actually up before logging into anything. They may want visibility into dozens or hundreds of remote sites from one place. In some cases, they may want an automated recovery process to take action before a human engineer ever gets involved.
That is a much different objective than simply providing remote access. Getting into the device is not the end goal. Getting the business back online is.
The Network Has Become Part of the Business
This becomes even more important as networks continue moving outside of traditional data centers and into thousands of distributed locations.
Think about what a retail store looks like today. It may have point-of-sale systems, payment processing, Wi-Fi, security cameras, digital signage, inventory systems, cloud applications, IoT devices, switches, routers, firewalls, and SD-WAN infrastructure.
Twenty years ago, we might have called that a branch office. Today, in many ways, it is a small data center.
The same thing is happening in healthcare clinics, bank branches, warehouses, manufacturing facilities, restaurants, transportation hubs, and edge environments.
And most of those locations do not have a network engineer sitting in the back room waiting for something to break. That changes the economics of an outage pretty quickly.
The issue is no longer simply that a router is down. The issue is that transactions may stop, applications may become unavailable, employees may not be able to work, customers may not be able to complete purchases, or an entire location may effectively go offline.
As businesses become more dependent on connectivity, the ability to recover the network remotely becomes just as important as the ability to keep it running.
Recovery Should Be Part of the Design
I think one of the mistakes we make is treating recovery as something we figure out after an outage occurs. It should really be part of the original architecture.
When designing a network, we should be asking what happens if the primary network becomes unreachable. Is there another way into the infrastructure that does not depend on the network we are trying to repair? Can we reboot a device remotely? Can we see what is happening at the location before we start troubleshooting? Can we manage hundreds of locations centrally? Can certain recovery actions eventually be automated?
Those questions become especially important when the infrastructure is remote or when downtime is expensive.
The best time to figure out how you are going to recover a network is not at 2AM when the network is already down.
The Next Step in Network Resilience
The networking industry has done an incredible job improving availability over the years. Networks are more redundant, more intelligent, and more automated than they have ever been. At the same time, the environment around those networks is changing.
We are managing more remote locations. Infrastructure is becoming more distributed. AI data centers are increasing the value of the infrastructure sitting behind the network. Engineering teams are expected to manage more devices with fewer people. And businesses have very little patience for extended outages.
I think all of that will force us to broaden the way we think about resilience. It cannot only be about preventing failure. It also has to be about what happens after failure.
Can we see the problem? Can we reach the infrastructure? Can we diagnose what happened? Can we take action remotely? And how quickly can we restore normal operations?
Redundancy will always be important. It reduces the likelihood that a single failure becomes an outage. But no amount of redundancy guarantees that something will never go wrong.
Eventually a circuit will fail, a device will stop responding, an upgrade will create a problem, or someone will make a configuration mistake.
When that happens, the organization with the better recovery architecture is usually going to be in a much better position.
We have spent decades building networks designed to stay up. I think the next step is building networks designed to get back up when they don’t.


