Failover configurations can look complete on paper and still fail silently. Small parameter mismatches between an application like SAP and the underlying replication layer go undetected through routine monitoring. They surface only when a real outage forces a failover and the cluster does not behave as expected.
In this interview on TFiR, Alexus Gore, Customer Experience Software Engineer at SIOS Technology, covers the most common SAP DataKeeper configuration gaps, explains why teams avoid failover testing and how to remove that risk, and walks through the two tools every team needs before touching production: a QA cluster and a runbook.
Guest: Alexus Gore, Customer Experience Software Engineer at SIOS Technology
Show: TFiR
Here is what every platform engineer and high-availability architect needs to know.
Technical Deep Dive
Q: What are the most common failover configuration issues that only appear during an actual outage?
Alexus Gore, Customer Experience Software Engineer at SIOS Technology, says the most frequent issue is missed application-level parameters when SAP is configured alongside DataKeeper. These parameters are small enough to escape notice during normal operation and are not flagged unless a proactive health check is performed. Because they sit below the threshold of routine log reviews, they remain invisible until a failover event forces the cluster into a state it was never correctly configured to handle.
“The most common issue is definitely not having your application set up in respect to the failover solution that you’re using.”
Alexus Gore, Customer Experience Software Engineer, SIOS Technology
Q: How do you test failover without disrupting production?
Gore recommends two concrete steps before any production testing. First, build a QA cluster: a one-to-one exact copy of the production cluster used exclusively for testing changes before they reach production. Second, create a runbook that records every step taken during testing, what was expected, what actually happened, and any unexpected disruptions. These two tools together let teams validate behavior safely and trace any deviation back to its source.
“To avoid any unnecessary risks, create a QA cluster and create a runbook to keep track of everything that you did.”
Alexus Gore, Customer Experience Software Engineer, SIOS Technology
Q: What should be on a failover testing checklist so nothing critical gets missed?
Gore identifies the runbook as the single most important item on any failover testing checklist. A runbook provides a documented record of every step, outcome, and deviation during a test. Without it, teams responding to an incident cannot trace what was done, what changed, or what should have prevented the failure. Gore notes that customers frequently cannot answer basic incident questions because no documentation existed at the time of testing.
“Having those set of steps in place, having documented information, is the most crucial when it comes to failover validation and failover testing.”
Alexus Gore, Customer Experience Software Engineer, SIOS Technology
Q: Do SIOS health checks catch configuration drift before a failure occurs?
Gore confirms that SIOS updates DataKeeper to detect and surface configuration issues before they cause failures. Log reviews and health checks are the mechanism: when a health check is performed, parameter mismatches that would otherwise stay hidden become visible. Teams that skip regular health checks lose this early-warning layer and shift the detection point to the outage itself.
“It can get missed, but it will get caught by just checking the logs because we do update our product to make sure that we’re catching these things beforehand.”
Alexus Gore, Customer Experience Software Engineer, SIOS Technology
Resources & Documentation
- SIOS Technology, vendor documentation and support resources for SIOS DataKeeper and SAP high-availability clustering
***
👇 Click to Read Full Raw Transcript
Swapnil Bhartiya: You talk to a lot of customers, you talk to your teams and I’m sure you have heard a lot of stories. What are the most common failover configuration issues that you have seen? The ones that don’t show up until an actual outage happens?
Alexus Gore: So most common, I would say happens when, sometimes when customers are using, let’s say, SAP applications with DataKey. And a lot of the times there are, we have documentation that cover the steps that are required to configure your SAP application properly to be used with Datakeeper. But sometimes the things that are like missed in the cracks may be like small parameters that are not updated to be used with datakeeper and they’re not usually caught until an outage occurs because it’s something that’s so small. This customer, we haven’t done a health check with data, then you’re not going to really see it until an outage occurs that there may be a problem happening. And a lot of times it does get caught. It can get missed, but it will get caught by just checking the logs because we do update our product to make sure that we’re catching these things beforehand. But yeah, I think the most common issue is definitely not having your application set up in respect to the failover solution that you’re using.
Swapnil Bhartiya: No, the fact is that a lot of IT teams, they hesitate to test failover because they worry about disrupting production. And that’s also a valid concern. How should organizations approach failover testing so that they build that confidence that things will not break? They get to test so they are not taking unnecessary risk.
Alexus Gore: So two things. To avoid any unnecessary risks, you would, ideally you would want to create a QA cluster beforehand before you even go into production. And a QA cluster is typically going to be a one to one exact copy of a production cluster that’s primarily going to be used for testing. So that way anything that you change or that you plan to change for your production cluster, you can do it in the QA cluster first. Second thing is creating a runbook. A runbook is usually going to have the steps that you ran during your testing. What happened, what occurred, what you expected to happen, maybe why it even didn’t happen. So this is good to keep track of everything that happened up to the moment where maybe a disruption occurred that you really didn’t expect to happen there. But yeah, to avoid any unnecessary risks, create a QA cluster and create a runbook to keep track of everything that you did.
Swapnil Bhartiya: Can you also share your thoughts? You know, what are some of the really crucial things when it comes to testing. You know, this checklist that they should not miss. One of those.
Alexus Gore: The biggest thing that you should not miss when it comes to testing is, first of all, with creating a runbook, which is going to have the list of steps that you ran during your testing to kind of like keep track of everything that happened. You want to make sure that everything that you tested was documented in the event that a situation happens, so you can kind of trace back your steps to see what could have been in place to prevent something from happening again. Runbooks, I think alongside QA clusters, runbooks, I feel, are the most crucial thing. A lot of the times we’ll have events or situations that have happened and we’re asking customers, oh, like, what happened when this occurred? And it’s like, I don’t really know what happened when this occurred, or I don’t really remember what I did. But having those set of steps in place, having documented information, is the most crucial when it comes to failover validation and failover testing.





