High availability environments fail silently. Configurations drift from design intent, patching cycles introduce unintended failover triggers, and pre-production environments that do not mirror production give false confidence. By the time an issue surfaces, it is already a production incident.
In this interview on TFiR, Matthew Pollard, Customer Experience Software Engineer at SIOS Technology, walks through the proactive communication strategies, testing disciplines, and patching considerations that IT admins need to keep HA environments stable and reliable.
Guest: Matthew Pollard, Customer Experience Software Engineer at SIOS Technology
Show: TFiR
Here is what every IT admin and platform engineer managing high availability infrastructure needs to know.
Technical Deep Dive
Q: How does SIOS Technology help IT admins address common high availability challenges?
Matthew Pollard, Customer Experience Software Engineer at SIOS Technology, says the company’s primary focus is proactive communication with partners and users. SIOS follows up after reported issues to confirm remediation steps have been completed and checks whether the current configuration is genuinely meeting the customer’s needs. The company also offers services to help customers identify and fix problems before they cause outages, including system checks and testing support.
“If I had to sum up our strategy, it’s just being proactive, open and communicative with our base.” — Matthew Pollard, Customer Experience Software Engineer, SIOS Technology
Q: What is the most important advice for IT admins looking to strengthen their HA setup?
Pollard’s top recommendation is to audit coverage first and confirm that everything users and teams depend on is actually protected under the HA configuration. After coverage is confirmed, rigorous and repeated testing is the most critical discipline. Testing should be as realistic as possible, which means the pre-production or QA environment must closely mirror production to produce meaningful results.
“Test, test, test, test, test it all. Test it as realistically as you can.” — Matthew Pollard, Customer Experience Software Engineer, SIOS Technology
Q: Why does patching and updating create risk in a high availability environment?
Pollard explains that HA solutions manage multiple interdependent components, and an unplanned or incorrectly sequenced patch can cause the HA solution itself to trigger a failover or generate alerts at exactly the wrong moment. When that happens, the failover or alert can interfere with the components currently being patched, compounding the risk. Admins should check with their vendors and follow the recommended patching steps precisely to avoid this cascading failure mode.
“If your patching causes the HA solution to trigger a failover or some kind of alert when it shouldn’t, that can affect the components that you’re patching as well.” — Matthew Pollard, Customer Experience Software Engineer, SIOS Technology
Q: How should IT admins approach planning for updates and patches in an HA environment?
Pollard stresses that proactive planning is essential. Every update or patch cycle should account for how each component interacts with the broader HA solution, not just the component being changed. Admins should verify vendor-recommended procedures before executing any change and treat thoroughness as a non-negotiable standard rather than an optional step.
“Be proactive with your planning. Make sure that when you have to update, when you have to patch, that everything is considered.” — Matthew Pollard, Customer Experience Software Engineer, SIOS Technology
Resources & Documentation
- SIOS Technology, high availability clustering software and services for business-critical applications
***
👇 Click to Read Full Raw Transcript
Swapnil Bhartiya: How does SIOS Technology approach these common high availability challenges that we just talked about to make life easier for IT admins?
Matthew Pollard: Right. So the number one priority is always communication. It’s always getting in touch with our partners, with our users and making sure that what they’ve configured is actually meeting their needs. It’s following up after they’ve noticed some kind of issue and saying, have you remediated it? Have you taken these steps that we’ve identified to make sure it won’t happen again? Is there any way we can help you make sure that your solution is meeting your needs? Any testing? Do you need us to come in and check your systems? We offer all kinds of services to help customers identify any problems and remediate them before they actually cause any issues. So I would say that if I had to sum up our strategy, it’s just being proactive, open and communicative with our base.
Swapnil Bhartiya: And finally, what advice would you give to IT admins looking to strengthen their HA setup?
Matthew Pollard: If you’re looking to strengthen your HA setup, I would say go through and really make sure that everything your users need, everything your teams need is being protected. I’m going to say it again because it’s just so important. Test, test, test, test, test it all. Test it as realistically as you can. If you’re setting up some kind of QA or pre production environment, then make sure that that is as similar as it can be to the production environment because that’s how you get the most credibility out of your testing. And be proactive with your planning. Make sure that when you have to update, when you have to patch, that everything is considered as a nature of a HA solution, having to manage all these different components. If your patching causes the HA solution to trigger, you know, a failover or some kind of alert when it shouldn’t, that can affect the components that you’re patching as well. Check with your vendors to make sure that you’re following their appropriate and recommended steps. It’s just it requires you to be very thorough. So that’s probably the best advice that I can give. Be thorough and be proactive.





