One could say AI is transforming how companies build their software, but it is also making them take a good, hard look at the infrastructure behind it. When firms are training large language models or running inference workloads and churning through huge datasets, they need the speed and scale that their legacy infrastructure can’t provide.
The numbers back this up IDC has global AI infrastructure spending on track to top $200 billion in 2028, and Gartner figures that over the next few years, more than three-quarters of all data an enterprise produces will be handled well beyond the confines of a conventional data centre. In short, for those in infrastructure, automation has moved beyond being a matter of efficiency to become a business necessity.
Old Playbook Doesn’t Work Anymore
Firms have to admit that their infrastructure put in place over the years was never really meant to handle the demands of an AI-centric workload. In the old days, when AI was still something of an experiment, one could get away with manually provisioning GPU clusters or keeping a close eye on a long training run. But that is not how things are done now. An enterprise will run hundreds of AI pipelines and inference services, processing millions of requests a day, with data engineering workflows that can scale from terabytes to petabytes within hours.
Then there is the financial side. Gartner has the numbers to back up why automation is so critical: companies that have automated their AI infrastructure see provisioning times drop by up to 73% and nearly 40% less compute waste. Idle GPUs and delayed requests are nothing but a drain on firms’ budget and productivity. With AI being adopted at such a pace, organisations can no longer view an automation-first approach as a choice; it is the only way to operate sustainably.
What “Automation-First” Really Means?
There’s more to the notion of automation-first than just setting a handful of mundane tasks on “autopilot.” What IT really means is an infrastructure that lets automation run its course and intervenes only when there is a problem. Think about Infrastructure-as-Code (IaC). With tools such as Terraform, Pulumi or Crossplane, one can reproduce everything from version-controlled code. In that way, data scientists can have their environments provisioned by an automated workflow in a matter of minutes rather than sitting around for days waiting on approvals.
But deployment is only part of it. Firms want intelligent autoscaling tied to GPU usage, and the ability to schedule workloads based on what matters to the business. If a training job gets cut off, the system should handle the checkpoint recovery on its own. It is no small thing either; Flexera’s 2025 State of the Cloud Report puts it at 84% of enterprises who put cloud cost optimisation at the top of their list. With AI costs rising so quickly, having that kind of infrastructure automation is essential to keep them in check.
GPU Problem Nobody Talks About Enough
In the world of enterprise AI, GPUs are now the single most important resource you have, yet also the one most likely to be left on the shelf. One would think supply and demand would ensure every bit of hardware is put to work, but in practice, many companies are not getting the most out of what they have.
The numbers back this up. Weights & Biases has conducted research showing that for the typical AI team, GPU utilisation hovers in the 35-45% range; in other words, well over half of your compute capacity is doing nothing. When one considers a top-tier AI accelerator running tens of thousands of dollars, any gains in how you use them can put millions back in a large enterprise’s pocket over the course of a year.
The way to close that gap is with automation. Firms can let cluster autoscalers quickly take down idle nodes or have Kubernetes schedulers do the heavy lifting of bin-packing workloads. There are profiling tools that will spot a wasteful training job before it gets underway. The best AI platform outfits are already treating GPU utilisation with the same kind of rigour as they do system uptime or application latency.
Observability Is the Foundation, Not an Afterthought
One can only get as much out of automation as one puts into it. And still, many organisations will invest in an automation platform without giving enough thought to the observability layer needed for it to be truly intelligent. The old way of looking at infrastructure metrics like CPU or network utilisation no longer cuts it when one is dealing with AI. You need to see what is happening with GPU memory, tensor throughput, storage bandwidth and the cost attribution at the workload level, to name a few. Lacking that kind of visibility means your automation cannot make good decisions.
There is a move in the industry towards this. OpenTelemetry is now the go-to for cloud-native observability, and there are tools like Grafana, Arize, Weights & Biases and NVIDIA DCGM for more focused monitoring of AI workloads. The numbers back it up: Splunk’s State of Observability report shows that companies with well-established observability practices put out fires much more quickly and keep operational costs down compared to those with more rudimentary monitoring.
Conclusion: Start Small, But Start Now
You will not find an organisation on the cusp of automating everything in one fell swoop and expecting to be done with it. For those just starting out, the aim is to pick a single process that is both repetitive and high-impact, then put in place the means to automate it end-to-end. Whether that means having idle GPU clusters shut down on their own, building a self-service portal for AI staff, or setting up cost monitoring for training workloads, the point is to do it end-to-end. Look at the ones who are ahead of the curve in AI infrastructure now, and you won’t always see the deepest pockets. What they have in common is an engineering culture where automation is second nature. An engineer there does not wonder if something can be automated, but rather why it was not done in the first place.

Authored by: Sameer Kadam, Vice President, Infrastructure Engineering

















