GPU, RAM, and storage costs are rising faster than infrastructure budgets can absorb. Enterprises that built procurement strategies around a single preferred vendor or assumed cloud rental would insulate them from hardware economics are now facing supply shortfalls, price spikes, and cost multipliers they did not plan for.
In this interview on TFiR, Rob Hirschfeld, CEO at RackN, covers how enterprises are adapting procurement strategy, why heterogeneous hardware environments are becoming the default, and why owning hardware is increasingly the only way to control AI infrastructure costs.
Guest: Rob Hirschfeld, CEO at RackN
Show: TFiR
Here is what every infrastructure architect and platform engineer needs to know.
Technical Deep Dive
Q: How is GPU and RAM scarcity affecting enterprise AI infrastructure procurement right now?
Rob Hirschfeld, CEO at RackN, explains that scarcity is affecting the entire industry because the cost of RAM and storage, alongside GPU and CPU costs, is rising sharply. Enterprises are locking in server purchases as early as possible to secure pricing and guarantee RAM shipments. The pressure is severe enough that waiting on procurement is no longer a viable strategy.
“The cost of RAM and storage, in addition to the cost of GPUs and CPUs, is going up to the extent where what we’re seeing people try to do is lock in server purchases as soon as possible so they can lock in prices and get shipments on RAM even today.” — Rob Hirschfeld, CEO, RackN
Q: Why are single-vendor hardware strategies failing under current supply chain conditions?
Hirschfeld notes that preferred vendors may be unable to supply hardware at all, or their costs may become prohibitive for teams committed to staying on a single platform. Being a Dell shop, a Cisco shop, or an HP shop is no longer a reliable procurement identity. Supply chain volatility means teams will get what they can get, not what they prefer.
“You need to realize you are not going to be able to have a preferred vendor. If you’re used to buying from one vendor, they might not be able to supply you or their costs might get prohibitive.” — Rob Hirschfeld, CEO, RackN
Q: What does a heterogeneous hardware strategy look like in practice for AI infrastructure?
Hirschfeld describes three behaviors he sees from customers adapting to scarcity: mixing vendors across the environment based on availability, compromising on system configuration to accommodate what ships, and sourcing RAM independently or planning to add it post-deployment. Teams are designing for heterogeneity upfront rather than treating it as an edge case. Even teams committed to one vendor are ending up with mixed environments.
“They are being heterogeneous in design upfront. They’re recognizing they’re either going to have to switch vendors and move, intermix different vendors depending on their supply chain, or compromise on how they configure the systems.” — Rob Hirschfeld, CEO, RackN
Q: How are enterprises extending server lifecycle to manage hardware scarcity?
Hirschfeld points to two related strategies: keeping systems in service longer through ongoing patching, updating, and revision, and sourcing hardware from the secondary market. As frontier CPUs and GPUs come off training workloads, they become available for inference at lower cost. This makes secondary market servers a practical option for teams that cannot secure new hardware at acceptable prices.
“They are planning to keep systems in service for longer, which means patching and updating and revising them, or looking into the secondary market for these servers as the frontier CPUs and GPUs come off market.” — Rob Hirschfeld, CEO, RackN
Q: Why is renting GPU capacity from cloud providers financially risky for AI workloads?
Hirschfeld is direct on this point: every hardware cost in the market gets passed through to tenants with a multiplier. Avoiding hardware ownership does not avoid hardware costs. Those costs surface in inferencing spend, cloud component fees, and any other line item where a service provider sits between the team and the physical infrastructure. Owning hardware is, in his framing, the only way to actually control that cost.
“Owning the hardware and controlling the cost of that hardware is absolutely essential because all of these costs are being passed on with multipliers from the service providers, and that is a very, very serious concern.” — Rob Hirschfeld, CEO, RackN
Q: Are ARM-based servers a viable option for enterprise AI inference workloads?
Hirschfeld raises ARM servers based on Nvidia chips as a real possibility teams should be prepared to evaluate. The underlying message is that hardware flexibility is no longer optional. Teams that have defined their infrastructure identity around a specific architecture or vendor family will need to expand what they consider acceptable hardware for AI inference.
“It even might be ARM servers based on Nvidia chips to do this work. You need to be prepared for a higher degree of flexibility in what type of gear that you’re onboarding.” — Rob Hirschfeld, CEO, RackN
Resources & Documentation
- RackN, infrastructure automation platform for managing heterogeneous bare metal and edge environments
***
👇 Click to Read Full Raw Transcript
Swapnil Bhartiya: Now, one of the biggest challenges these days is getting AI gear. It is one of the hardest practical challenges right now. How are customers dealing with GPU and RAM scarcity?
Rob Hirschfeld: Yeah, this is something that’s affecting the industry as a whole because the cost of RAM and storage, in addition to the cost of GPUs and CPUs is going up to the extent where what we’re seeing people try to do is lock in server purchases as soon as possible so they can lock in prices and get shipments on RAM even today. And so this is one of those places where you need to realize you are not going to be able to have a preferred vendor, that if you’re used to buying from one vendor, they might not be able to supply you or their costs might get prohibitive for you to continue to stay on that platform of choice. What we see customers doing is they are being heterogeneous in design upfront. So they’re recognizing they’re either going to have to switch vendors and move, be able to intermix different, different vendors depending on their supply chain. They’re going to have to compromise on how they configure the systems and then have heterogeneous environments, even if they’re sticking with one vendor or they are sourcing RAM themselves or potentially planning to add RAM later, assuming, let’s hope it frees up in the market. And they are planning to keep systems in service for longer, which means patching and updating and revising them, or looking into the secondary market for these servers as the frontier CPUs and GPUs come off market. They’re very usable for inference and they’ll be available for you from these training labs. So you need to be looking very differently at the hardware infrastructure that you might have said, oh, I’m a Dell shop or a Cisco shop or an HP shop, and been thinking that would save you. The reality today is that you’re going to get what you get. It even might be ARMS servers based on Nvidia chips to do this work. And you need to be prepared for a higher degree of flexibility in what type of gear that you’re onboarding. And if you’re thinking that’s scary, I just want to never own hardware again. The challenge of this market is that renting your hardware or getting it from a service provider, you’re going to be paying that markup multiple times over. And so while it might be scary to pay more for hardware, owning the hardware and controlling the cost of that hardware is absolutely essential because all of these costs are being passed on with multipliers from the service providers, and that is a very, very serious concern. If you’re trying to manage your budget, it’s going to show up in your inferencing costs, it’s going to show up in your cloud and other component costs, it’s going to show up in basically any way you turn trying to avoid having hardware, the hardware costs are still going to get passed down to you.





