Having a failover solution configured does not mean it will behave correctly when a network outage or system disruption occurs. The gap between configuration and verified readiness is where production incidents happen. Without deliberate testing and log-level visibility, unknowns remain in the environment until the worst possible moment.
In this interview on TFiR, Alexus Gore, Customer Experience Software Engineer at SIOS Technology, walks through how to determine whether a failover environment is genuinely ready, what logs to check first under time pressure, and how system and failover logs surface cluster communication failures before they become outages.
Guest: Alexus Gore, Customer Experience Software Engineer at SIOS Technology
Show: TFiR
Technical Deep Dive
Q: What is the biggest difference between having failover configured and knowing it will actually work during an outage?
Alexus Gore, Customer Experience Software Engineer at SIOS Technology, explains that the critical differentiator is testing. A configured failover environment may still contain unknowns about how the system will react to a network outage or any other disruption. If those unknowns exist, the failover solution cannot be considered reliable. The goal is to reach a state where the outcome of any failure scenario is fully predictable, because untested assumptions are the direct cause of failed recoveries.
“If there are any unknowns into whether or not that failover solution is going to react appropriately, that is typically not the best case scenario.” — Alexus Gore, Customer Experience Software Engineer, SIOS Technology
Q: What should you check first when you have five minutes to assess the health of a production failover environment?
Gore recommends starting with logs. Logs are the fastest path to understanding the current state of a failover environment and will surface most active or emerging problems quickly. Both failover solution logs and system logs should be reviewed, as each covers a different layer of the environment. Together they provide a complete picture of cluster health and system communication status.
“Logs are generally going to tell you everything you need to know that’s happening in your failover environment.” — Alexus Gore, Customer Experience Software Engineer, SIOS Technology
Q: What do failover solution logs specifically tell you about your environment?
Failover solution logs show how the failover software itself is handling the environment. They surface network issues and communication problems occurring within the cluster. This makes them the primary diagnostic layer for understanding whether the failover solution is operating as expected and whether any cluster-level communication breakdown is underway.
“Your failover solution logs are typically going to cover what’s going on with how your failover solution is handling your environment.” — Alexus Gore, Customer Experience Software Engineer, SIOS Technology
Q: What do system logs reveal that failover solution logs do not?
System logs provide visibility into issues at the operating system level, including network problems and communication failures between the system and the failover solution. Where failover solution logs focus on cluster behavior, system logs can identify misconfigurations or degraded communication that originates outside the failover layer. Reviewing both in parallel gives a full diagnostic picture during a rapid health check.
“With your system logs, you can also determine if there are any issues with your system having any network problems or mishaps going on between your system and your failover solution.” — Alexus Gore, Customer Experience Software Engineer, SIOS Technology
Resources & Documentation
- SIOS Technology, high availability and disaster recovery software for business-critical environments
***
👇 Click to Read Full Raw Transcript
Swapnil Bhartiya: Most organizations, they assume that they are protected just because they have failover solutions in place. And from their perspective, there is nothing wrong. That’s what they should assume. But can you talk about what is the biggest difference between having failover configured and actually knowing that it will work when they actually need it?
Alexus Gore: So the biggest difference in determining how your failover environment is going to work when you need it is largely based in the testing that is that has been completed for it. With testing, you need to kind of like check the ins and outs with how your environment is going to react in a situation like should a network outage occur, and if there is like an unknown answer into whether or not you know you have your failover configured into whether or not that failover solution, you know it’s going to react appropriately. In the event that a failover outage occurs, if there are any unknowns there, then that is typically not the best case scenario. You want to be able to know what’s going to happen, how your solution is going to react in the event that a network outage occurs, in the event that any type of disruptance occurs, ideally,
Swapnil Bhartiya: and let’s assume that you have just five minutes to assess the health of a production failover environment. What are the first few things you would check?
Alexus Gore: First few things I check, I think mainly I’d start with the logs. Logs are generally going to tell you everything you need to know that’s happening in your failover environment.
Swapnil Bhartiya: You have your failover solution
Alexus Gore: logs, and your failover solution logs are typically going to cover what’s going on with how your failover solution is handling your environment. And you can also check your system logs. Your system logs can usually tell you, I mean, both will usually tell you if there are any network issues or, or communication issues that are happening within a cluster or in your environment. But with your system logs, you can also kind of determine if there are any issues with your system having any network problems or just like lack of, I guess, communication or just like mishaps going on between your system and your failover solution.





