Enterprises buying GPU hardware for AI workloads are discovering that standing up a cluster manually is only the beginning of the problem. The resulting environments are bespoke, untestable, and impossible to patch or recreate without starting over. Meanwhile, cloud alternatives lack the GPU capacity, cost controls, and governance that production AI requires. The gap between a working lab cluster and a production-grade, automatable infrastructure platform is where most AI initiatives stall indefinitely.
In this interview on TFiR, Rob Hirschfeld, CEO at RackN, covers how the Digital Rebar platform automates the bare metal layer to make hardware onboarding repeatable and predictable across heterogeneous server environments, and why starting with production-grade operations from day one is the decisive factor in moving AI workloads out of perpetual POC.
Guest: Rob Hirschfeld, CEO at RackN
Show: TFiR
Here is what every platform engineer and infrastructure decision-maker needs to know.
Technical Deep Dive
Q: What infrastructure decisions should executives make first when moving AI past experimentation?
Rob Hirschfeld, CEO at RackN, frames the core decision as one of ownership and control: executives need to determine how they will own the infrastructure layer that runs their AI processing, guarantee its availability, and manage token costs as capacity demands grow exponentially. Hirschfeld argues that reaching for bare metal and on-premises infrastructure is the rational response to cloud GPU scarcity, cost pressure, and governance requirements. The prerequisite is having an automation platform in place before hardware is purchased, not after.
“Most companies want their token capacity to grow exponentially over the next couple years. How do I own this core part of my business processing, how do I guarantee its availability, how do I make sure that I can afford the tokens I do want to spend?” — Rob Hirschfeld, CEO, RackN
Q: What does RackN’s Digital Rebar platform actually do at the bare metal layer?
Digital Rebar encodes proven operational practices into the provisioning pipeline so that any hardware plugged into the system moves through a repeatable onboarding sequence: imaging, cluster joining, patching, firmware updates, and network configuration. Hirschfeld emphasizes that the platform handles heterogeneous environments by design, with customers running ARM, Intel, and AMD servers within the same cluster. The outcome is that teams can bring Kubernetes clusters online in one to two weeks rather than the months or years that manual provisioning typically requires.
“What RackN does through Digital Rebar is make the bare metal layer specific, predictable, and repeatable. We’ve embedded battle-proven operational practices into the platform.” — Rob Hirschfeld, CEO, RackN
Q: Why are enterprises moving AI workloads away from cloud and back to on-premises bare metal?
Hirschfeld identifies three converging pressures: cloud costs are unsustainable at AI scale, GPU availability in public cloud does not meet demand, and enterprises require data governance controls that cloud architectures do not easily provide. He draws a direct parallel to the original cloud migration: organizations that gave up on-premises infrastructure to escape bespoke, fragile, expensive data centers are now finding that cloud introduces a different set of constraints that are equally limiting for AI workloads.
“Cloud is expensive, doesn’t have all the GPU resources they need, doesn’t have the controls and governance that they want, or they don’t want all that data flowing all over the place. They’re looking at how do I pull it back?” — Rob Hirschfeld, CEO, RackN
Q: What is the biggest mistake teams make when setting up AI infrastructure?
Hirschfeld identifies hand-building clusters as the critical failure mode. Teams purchase expensive GPU hardware, configure it manually, and produce an environment that is unique, undocumented in any reproducible way, and impossible to patch, upgrade, or reset without breaking it. The downstream consequence is that those environments can never be promoted to production because there is no automated path to bring a new node into conformance with the existing cluster state.
“A lot of people buy expensive AI gear and then they set it up by hand and they end up with a bespoke environment that they don’t know how to recreate, they don’t know how to upgrade, they don’t know how to patch.” — Rob Hirschfeld, CEO, RackN
Q: Can AI tools be used to manage bare metal operations and infrastructure provisioning?
Hirschfeld is direct: AI tools are unreliable for bare metal operations because there is insufficient training data covering that class of work. RackN has observed customers attempting to use AI for operational tasks and consistently producing bricked systems, poorly optimized configurations, or environments that cannot be brought to a known-good state. The absence of a strong training corpus means AI-generated operational instructions carry high risk in physical infrastructure contexts.
“If you ask an AI to help you with operations, they are not reliable sources. There’s no training data for the type of work that we do. We watch customers try to ask AI to do operations for them and they end up with bricked systems.” — Rob Hirschfeld, CEO, RackN
Q: How does automated reset and refresh work in a production bare metal environment?
Within Digital Rebar, a reset operation re-runs the full provisioning pipeline against an existing node, returning it to a known-good state with the current firmware, OS image, and network configuration applied. Hirschfeld describes this as critical for maintaining cluster conformance as hardware cycles in and out of service: any server rejoining the cluster automatically receives the latest patches and firmware updates before being allowed to carry workloads. The reset capability also accelerates experimentation because teams can rapidly iterate on configuration changes without manual remediation steps.
“Once it’s built, they can just push a button, reset and bring it back and keep it in conformance. As the systems come in and out, they’re automatically getting patched and updated, they’re getting the latest firmware patches, they’re getting on the networks.” — Rob Hirschfeld, CEO, RackN
Q: How does Digital Rebar handle advanced networking hardware like DPUs and SmartNICs?
Hirschfeld acknowledges that DPUs and SmartNICs represent a layer of complexity on top of already heterogeneous server environments, and that the platform’s provisioning pipeline addresses advanced networking configuration as part of the standard onboarding sequence. He frames this as part of the broader value proposition: customers do not need to solve each layer of hardware complexity independently because the platform encodes those operational patterns, allowing teams to focus on the workloads running above the infrastructure layer.
“We haven’t even talked about the DPUs and the SmartNICs and things like that. It’s a very complex environment out there. Our customers are able to just use their infrastructure and then focus on the workloads on top of it.” — Rob Hirschfeld, CEO, RackN
Q: Why do skills learned in lab-phase AI infrastructure work not translate to production?
Hirschfeld draws the distinction between the complexity tolerance of a personal lab environment and the operational requirements of production: a server in a basement is acceptable to rebuild manually because the consequences of failure are low. Production AI infrastructure involves interdependent clusters, SLA commitments, firmware compatibility matrices, and network configurations that cannot be safely managed by the same ad-hoc methods used for learning. The skills and muscle memory built in a lab are specifically adapted to conditions that do not exist in production.
“The things you learn in that lab phase won’t translate. I promise you they do not translate into production. So the sooner you get to production-grade operations, the faster you’re going to get to end results.” — Rob Hirschfeld, CEO, RackN
Q: What does a perpetual POC look like and how do teams break out of it?
Hirschfeld describes a perpetual POC as a hand-built cluster that teams continue iterating on without any automated path to production conformance. Because the environment cannot be reliably reproduced, every proposed change requires manual remediation, and the cluster accumulates technical debt that makes promotion to production increasingly risky. His prescription is to begin with production-grade automation tooling from the first hardware deployment so that the POC environment is itself a reproducible artifact, not a unique snowflake.
“The biggest mistake that we see people doing is they get stuck in perpetual POCs that don’t have any hope of making it into production. You have to start with production grade.” — Rob Hirschfeld, CEO, RackN
Resources & Documentation
- RackN Digital Rebar, bare metal lifecycle management and automation platform for heterogeneous infrastructure
***
👇 Click to Read Full Raw Transcript
Swapnil Bhartiya: For executives who are trying to move past experimentation, what are the first infrastructure decisions they should make right now?
Rob Hirschfeld: I’m excited about this era that we’re entering because Rack N’s mission goes back a generation in helping people be able to run hardware and bare metal themselves, own their own infrastructure and make that easy and possible and scalable. And so what we’re really doing is helping companies who are looking at all of this AI complexity and cost and the need for governance and control and service level agreements around their AI infrastructure and, you know, evaluate how do I own this core part of my business processing, how do I guarantee its availability, how do I make sure that I can afford the tokens that I do want to spend? Because most companies want their token, their token capacity to grow exponentially over the next couple years. And what rackn does through our platform digital rebar is we make the bare metal layer specific, predictable and repeatable. We’ve embedded just amazing battle proven operational practices into the platform. So when you bring any hardware, and I mean we literally have customers bringing up ARM servers alongside Intel AMD servers, whatever servers you plug into the system, we’re able to onboard, take through a repeated process, get the platforms installed, join them into clusters, patch, update, network, all of those operational capabilities are baked into the platform. And it’s really important because in this world of incredibly heterogeneous, fast moving, expensive hardware, our customers are incredibly confident in their ability to onboard and use infrastructure in a way that, you know, we’re very proud of and I haven’t seen anywhere else in the industry. So a lot of people think about data centers and they think back to the 90s and it was incredibly bespoke and slow and expensive and very fragile. And you were always patching and you know, it was a real treadmill of operations. And they gave that up and moved to cloud. And now they’re finding that, you know, cloud is expensive, doesn’t have all the GPU resources they need, doesn’t have the controls and governance that they want, or they don’t want all that data flowing all over the place. And they’re looking at how do I pull it back? What we’ve been able to do with the platform is make it so that they can just show up and start working. We bring Kubernetes clusters up in a week or two weeks where our customers might have been spending months or even years trying to bring that hardware into conformance. And the beauty is once it’s built, they can just push a button, reset and bring it back and keep IT in conformance. As the systems come in and out, they’re automatically getting patched and updated, they’re getting the latest firmware patches, they’re getting on the networks, some advanced networking is going on. We haven’t even talked about the, the DPUs and the SmartNics and things like that. It’s a very complex environment out there. Our customers are able to just use their infrastructure and then focus on the workloads on top of it. That discipline, that out of the box capability and the ability to ingest whichever hardware they need to run whichever platforms they need on top of it is a game changer. And the confidence that you have that you can focus on getting your business done is really, really important. Because we’ve been talking about how challenging the environment is, how much things are changing, how much the expertise of getting this done is important. And I will promise you, if you ask an AI to help you with operations, they are not reliable sources. There’s no training data for the type of work that we do. And we watch customers try to ask AI to do operations for them and they end up with bricked systems, they end up with really poorly optimized systems, or they end up with just a huge mess. And so the ability to bypass that and just get straight into getting work done is absolutely essential. In this era, it’s very important to learn how to use the models and harnesses like we’ve discussed in detail. The thing that I would tell people is that be very careful with your experiments. What we have seen is that a lot of people buy expensive AI gear and then they set it up by hand and they end up with a bespoke environment that they don’t know how to recreate, they don’t know how to upgrade, they don’t know how to patch. You can’t afford to buy infrastructure and spend months and weeks fixing it, tuning it, doing all that stuff. Start with systems that have the automation. This is why we like to get involved with customers right in the start, is that we can help you get your clusters up and running. But more than that, get them automated so you can do a reset and a refresh and patch and control that actually lets you move these experiments faster. So don’t get confused that you have to turn every knob and build this from scratch. Like you might have the garage days or the basement ops themes where everybody had a server in their basement. That was a great way to learn, you know, last decade. Today you don’t have time for that. The systems are really complex and a lot of the things you learn in that lab phase won’t translate. I promise you do not translate into production. So the sooner you get to production grade operations, the faster you’re going to get to end results. And that’s the biggest mistake that we see people doing is they get stuck in perpetual pocs that don’t actually have end any hope of making it into production. You have to start with production grade.





