Infrastructure expert warns AI is creating a new data centre challenge

September 3, 2026 at 7:17 AM GMT+8

Synopsis: AI is rapidly transitioning from training environments to real-world inference deployments. This shift fundamentally alters data centre requirements. While Australia is uniquely positioned to capture ensuing infrastructure investments, success will depend on addressing critical bottlenecks: securing power, accelerating grid interconnection, scaling efficiently, and maintaining rigorous operational reliability.

AI is moving from the training campus to the real world and that shift will inevitably change what data centres need to deliver. Anahita Mouro [1], a digital infrastructure leader at Google specialising in field performance and process excellence, believes Australia is well positioned for the next wave of AI infrastructure investment, but says the winners will be determined not just by access to GPUs, but by the ability to secure power, connect to the grid, scale efficiently and operate increasingly complex infrastructure reliably.

For the past two years, much of the conversation around AI infrastructure has revolved around one thing: building ever-larger clusters to train ever-larger models. That era is not over. But according to Mouro, the industry’s next infrastructure challenge is already taking shape. And it looks very different from the training environments that have dominated the headlines.

“Two years ago, almost every infrastructure conversation was really a conversation about training; bigger clusters, more GPUs, faster interconnects, all to help build the next foundation model,” Mouro told W.Media. “What’s changed is that those models are now actually being used, at massive scale, in real time.” Gartner projects that by 2028, more than 80 percent of AI infrastructure spending will go toward inference, not training.

Australia’s position

That distinction is becoming increasingly important for Australia. Mouro, who leads a global organisation of senior technical programme managers and engineers responsible for the field performance and reliability of Google’s hyperscale infrastructure, will bring that perspective to the Sydney Cloud & Datacenter Convention in September.

Her keynote, “From Deployment to Decision: Building Infrastructure That Keeps Pace with Agentic AI,” will examine how the infrastructure built for today’s gigawatt-scale AI deployments will need to evolve as AI systems increasingly make decisions and execute tasks continuously.

The shift from training to inference is at the heart of that argument. “Training is a batch workload; if something fails, you checkpoint and restart. It is costly but it is not user-facing,” she says. “Inference is a production, user-facing workload that has to behave like any other critical web service: low latency, high availability, resilient under variable load.”

For Mouro, that changes the infrastructure conversation fundamentally. “So the conversation has matured. That’s why I think infrastructure, not the models themselves, is now the more interesting story in AI.”

The inference problem

The change becomes even more pronounced when AI systems become agentic. Today’s familiar model interaction is relatively straightforward: a user submits a request, the model generates a response and the interaction ends. Agentic systems introduce a chain of dependent actions: agents execute multi-step workflows, and usually wait on external systems before taking the next action.

“From a user perspective, all of those steps still need to feel like a single, responsive interaction,” Mouro says. That creates a different infrastructure challenge because the performance of the overall interaction is no longer determined by a single model call. “The dependent nature of agentic steps means that each has its own latency and reliability requirements, so the tail latency of the whole chain matters more than any individual call,” she says.

In other words, infrastructure has to cope not only with more AI computation, but with a much more complex sequence of computational events. “Agentic AI takes everything that already makes inference harder than training, [like] latency sensitivity, availability, variable load, and compounds it.”

Mouro describes managing the tail latency of these multi-step automated decisions as “a fundamentally different engineering problem than standard request-response inference.”

That has consequences well beyond the GPU itself.

From bigger clusters to different infrastructure

The growth trajectory is one reason Mouro believes the infrastructure implications of inference deserve absolutely more attention. “Gartner projects inference server spending growing around 42 percent annually through 2028, roughly double training’s growth rate,” she says.

The physical consequence of that growth is significant. “When you translate that growth into physical reality, we are talking about deploying gigawatt-scale data centres.”

In essence, a training run may produce a model once, but that model can subsequently be used for billions of inference operations. “One training run produces a model that then serves billions of inference operations, and as agentic workflows multiply the number of model calls per user task, that gap widens further,” she says.

That is beginning to change the design priorities for AI infrastructure. “Hardware procurement is splitting: latest-generation GPUs for training, and a more diverse mix of inference-optimised chips, older-generation GPUs, and purpose-built ASICs for serving,” she says.

Mouro adds that power and cooling are also shifting: from “provisioning for constant, predictable load to managing variable, bursty load dynamically.” And networks will need to respond to the same change, with infrastructure designed to support dynamic load balancing.

For Mouro, the conclusion is that inference infrastructure can no longer be treated simply as another deployment of compute capacity. “The organisations that get ahead here are the ones treating inference infrastructure with the same production-grade discipline they’d apply to any critical consumer-facing service, because that’s exactly what it is now.”

Infrastructure is the pacing factor

That discipline will become increasingly important because the physical infrastructure supporting AI is struggling to move at the same pace as the technology it houses. “Infrastructure is now the pacing factor,” Mouro says, adding that chip development has moved extraordinarily quickly, with successive generations delivering better performance per watt and an increasingly diverse range of accelerators.

“But hardware is only half the equation,” she stresses. Power availability, cooling capacity and grid interconnection timelines have not accelerated at the same rate as silicon. “A faster chip doesn’t help if you can’t get the power to run it or the cooling to keep it running.”

That, she says, makes the next stage of infrastructure innovation less about installing more powerful hardware and more about closing the gap between technological development and physical deployment. “Model and hardware innovation are outpacing the physical infrastructure that has to house and power them,” she says.

That means more efficient facility design, dynamic power management and smarter site selection will become increasingly important.

Australia should move quickly

Those issues have particular relevance to Australia. Mouro sees a strong underlying proposition for the country: stable governance, a sophisticated technical workforce, proximity to growing Asia-Pacific markets and significant momentum in renewable energy. “Australia’s opportunity is genuinely strong,” she says.

But she identifies two issues that need attention as the market scales: grid capacity and interconnection speed, particularly in regions suited to large-scale development, and the ability of planning and permitting processes to keep pace with the speed of the industry. “When you’re facing multi-year utility connection delays, you have to get creative,” she says.

Mouro points to approaches including behind-the-meter colocated energy generation and flexible load architectures as areas that deserve active consideration. “None of these are unique to Australia, and all solvable but the markets that solve them earliest normally capture a disproportionate share of the investment that follows,” she says.

That goes to the heart of Australia’s data centre opportunity. Having available renewable energy or suitable land is not sufficient if infrastructure cannot be connected and projects cannot progress on the required timeline.

“Power comes first, always: reliable, scalable access to energy, ideally with a credible path to low-carbon supply,” she says. “Grid interconnection speed matters just as much as generation capacity; a region with abundant power but a multi-year queue to connect it isn’t actually investable on the timelines this industry moves at.”

Connectivity and skills follow, with low-latency access becoming more important as inference becomes increasingly user-facing. Finally, she points to regulatory and political stability. “This is patient capital deployed over years or decades, and predictability in policy, permitting and land use matters as much as any single incentive on offer,” she adds.

The operational challenge

If power and infrastructure availability determine where AI facilities can be built, operational discipline increasingly determines whether they can perform as intended. Mouro says the sheer pace of deployment is itself one of the biggest challenges. “We’re deploying more capacity, faster, than at any point in this industry’s history,” she says.

That compresses the time available for commissioning, validation and operational readiness. “Doing that without cutting corners on reliability takes real discipline,” she says.

The second challenge is complexity. A hyperscale AI facility today brings together power systems, liquid cooling, networking fabrics and thousands of accelerators, all increasingly running workloads that shift dynamically between training and inference profiles, each with very different power and thermal signatures. “We are pushing all systems close to their limits in order to accommodate the velocity and scale,” she says.

That creates another challenge: consistency. Google operates infrastructure across regions with different climates, grids and regulatory environments. Standards and operational playbooks therefore need to be transferable without becoming so rigid that they cannot accommodate local conditions. “Getting that balance right, global consistency with local flexibility, is a constant, active piece of work, not something you solve once,” she says.

The same principle applies inside the facility. “TPUs, GPUs and power get the headlines, but a hyperscale AI facility is really hundreds of interdependent systems: cooling, power distribution, networking, telemetry and monitoring,” she says. “Failover and reliability is a property of how well those systems work together, not of any one of them in isolation.

“That’s why we design inference infrastructure the way you’d design any production-level service: redundant servers behind load balancers, automated health checks and failover, continuous monitoring and alerting, tested disaster recovery,” she says. “None of that is glamorous, but it’s what separates a facility that quietly hits 99.99% uptime from one that looks impressive on paper and fails under real load.”

From reactive to predictive operations

One of the more interesting changes Mouro sees is the growing role of digital twins and intelligent automation in managing that complexity. Digital twins can model power, thermal and network behaviour before changes are made to a physical facility. That allows operators to understand how a cooling modification or sudden load increase might propagate through a system without discovering the consequences during live operation.

“We’re moving from operations that are reactive to operations that are predictive: catching a thermal imbalance or a failover risk before it becomes an incident, not after,” she says. “That doesn’t replace operational discipline; it’s an extension of it.”

Runbooks, playbooks and postmortem cultures remain important. The difference is that digital twins and intelligent automation can allow that discipline to be applied across increasingly complex facilities without a proportional increase in headcount or risk.

The leadership shift

Mouro says the qualities infrastructure leaders need are changing as fast as the infrastructure itself. Chief among them is comfort with ambiguity. “The infrastructure requirements we’re planning against today will look different in eighteen months,” she says. “We’ve already seen that with the shift from training to inference, and agentic AI is the next iteration of it.”

Technical depth still matters, but increasingly, she says, the hardest problems sit at the intersections between disciplines rather than within any one of them. “I’d add operational humility,” she says. “This industry has a way of humbling anyone who assumes yesterday’s playbook still applies.”

The people behind the infrastructure

Another dimension to the AI infrastructure build-out that Mouro believes receives too little attention is the people required to deliver it. She recently joined the board of Nomad Futurist, an organisation focused on encouraging the next generation to pursue careers in digital infrastructure.

“This industry has a visibility problem,” she notes. Most people using AI have little understanding of the global infrastructure workforce required to make it possible. “That invisibility means we’re not attracting anywhere near the talent this build-out is going to need over the next decade.”

“I’ve spent my career in infrastructure, and some of the best people I’ve worked with, myself included, found our way here almost by accident, because nobody told us this was a career path when we were choosing one,” she says. “If that’s true for people who did end up here, it’s certainly true for a much larger group who never got the chance to consider it, particularly women and people from backgrounds this industry hasn’t traditionally reached.”

The opportunity, she argues, extends far beyond traditional data centre engineering. The current build-out requires power systems engineers, thermal engineers, networking specialists, mechanical and civil engineers and software and operations professionals.

But it also requires lawyers, healthcare and safety professionals, planners, environmental scientists, economists and communications specialists. “This is one of the broadest industries there is,” she stresses. “Almost any career path you can name has a place in this build-out,” she adds. “It’s also a build-out that’s going to touch every region, not just the traditional hubs.”

“The skills being developed right now, managing dense, variable, latency-sensitive systems at scale, will be valuable for decades, well beyond whatever the current hardware generation happens to be,” she says. “That’s a good career to bet on.”

The next infrastructure era

Her central message is not that training infrastructure is disappearing but it is that the industry needs to stop assuming that the infrastructure requirements of training will define the next decade.

“The infrastructure conversation has fundamentally shifted, and the organisations that recognise it earliest will have the advantage,” she points out. The next phase will increasingly be defined by inference and agentic systems: production infrastructure that is always running, exposed directly to users and expected to meet the reliability standards of other critical digital services.

“Different latency requirements, different availability standards, different hardware economics, different power and cooling profiles, different network patterns,” she says.

No single organisation is solving this alone, Mouro says. Power utilities, hardware vendors, hyperscalers, operators and regulators are all working on interlocking pieces of the same problem. Meanwhile, as AI infrastructure increasingly becomes a matter of national policy, she argues that good infrastructure policy needs to be grounded in operational reality.

“The next decade belongs to inference and, increasingly, to agentic systems: production-grade, user-facing, always-on infrastructure that has to be engineered with the same discipline as any critical service the world depends on,” she says. “If delegates [in Sydney] leave thinking about their infrastructure roadmap through that lens, production reliability and not just training capacity, the keynote will have done its job,” she says.

[1] Disclaimer: This article is based on a keynote presentation which will be delivered by Anahita Mouro in her personal capacity. The views, analyses, and opinions expressed in this piece are solely those of the author and do not represent the official stance, policies, or strategic direction of Google or Alphabet.

Anahita Mouro will deliver the international keynote “From Deployment to Decision: Building Infrastructure That Keeps Pace with Agentic AI” at the Sydney Cloud & Datacenter Convention at 9am on 17 September, as part of the Data Center Plenary. The presentation will bridge the realities of deploying infrastructure for gigawatt training campuses today with the future demands of agentic, always-on AI systems.

The Sydney Cloud & Datacenter Convention 2026 will take place at the ICC Sydney, Exhibition Centre, Hall 4 on 16–17 September 2026. To attend the event please visit the event website.