Loading Now

AI transformation across the infrastructure lifecycle: From supply chain to fleet operations

AI infrastructure is an interconnected system, where each component impacts the others. The choices we make during the development and design of our hardware shape how we acquire, deploy, and manage our technology across the board. Additionally, experiences we gain from the operation of this hardware can inform our future projects.

This continuous feedback is crucial as the landscape is always changing. Customer needs fluctuate, components can become scarce, new capacity is introduced, and the requirements for hardware adapt as well. To ensure our infrastructure remains dependable, it’s essential to keep learning and adjusting in response to these ongoing changes.

At Microsoft, the Azure Hardware Systems and Infrastructure team oversees this entire lifecycle. From system architecture and design to supply chain management, deployment, and operations across more than 80 regions and 500 datacenter campuses, this comprehensive perspective enables us to connect insights throughout the hardware lifecycle. By learning from one area of the system, we can enhance decision-making across different facets.

AI plays a significant role in accelerating this learning process. During our own AI transformation, we’ve harnessed AI tools to facilitate information sharing, helping teams stay informed about changes and take action swiftly, without losing the essential input of human judgement. The goal isn’t just to speed up tasks; it’s about constructing a system that learns from infrastructure design, procurement, and operations and applies those insights to future projects. This strategy is a key element of our wider AI transformation initiative.

Begin with the Task, Not the Technology

As we expanded our application of AI to enhance our cloud infrastructure, a vital lesson emerged: AI transformation should start with the work itself, rather than the technology. While speed is important, the bigger advantage lies in redesigning decision-making processes: ensuring the right information is available at the right moment, quickly understanding changes, and determining where human judgement is essential.

For instance, our cloud supply chain exemplifies this approach. Each month, demand-planning teams forecast Azure’s future infrastructure requirements, considering shifts in customer demand, regional specifics, installed capacity, and decommissioning activities. When plans shift, understanding the reason can often involve collating data from various systems, turning a straightforward inquiry into a time-consuming task. It would be easy to jump in and just implement an AI solution.

However, merely layering AI over a fragmented process can lead to faster disintegration of that fragmentation. Prior to implementing AI, our teams focused on mapping and streamlining workflows, establishing a central data foundation that ensured quality, governance, and access controls, and pinpointed where human accountability was crucial. We refer to this strategy as “Lean before AI.”

By starting with complete processes and applying an AI-driven mindset, we can initiate new working methods that promote parallel execution over sequential handoffs, fostering integrated and collaborative workflows.

From Days of Research to Quick Decisions

With our foundational changes in place, we approached demand planning with a new perspective. A multi-agent workflow can analyse signals like shifts in the installed base, regional demand, and changes due to decommissioning. This analysis helps planners discern what has changed, where it happened, and what caused the shift. Tasks that once took five to seven business days can now often be completed in mere hours, sometimes even in under 20 minutes. Over five monthly planning cycles, our demand-planning team experienced nearly a 50% reduction in manual tasks, with cycle times reduced by as much as 75% in selected workflows.

This successful pattern is emerging across various functions, including planning, product data management, sourcing, fulfilment, logistics, and operational workflows. By employing specialised agents, teams dedicate less time to gathering and reconciling information and more time to applying their professional expertise.

In the fulfilment area, for example, understanding why rack delivery is hindered can previously involve pulling data manually from multiple sources. Now, an intelligent assistant collates information about the blockers and checks for compatible or incompatible supplies, reducing investigation time by up to 55%. Similarly, in logistics, an AI-powered tool integrates data from air, land, and sea options, enabling teams to evaluate speed, cost, and carbon trade-offs while also predicting emissions.

Creating the Learning Loop

While these individual applications are significant, the true opportunity lies in connecting them. Our cloud supply chain team is transitioning towards comprehensive multi-agent workflows for tasks like bill-of-materials generation, capacity delivery, spare-parts management, and sales and operations execution. This initiative signifies a larger evolution in our organisation: moving from isolated trials to a more resilient, intelligent operating system that prioritises human judgement.

The main benefit of this approach isn’t just speed. Planners can now begin with interconnected evidence rather than spending days piecing it together. This shift allows them ample time to understand the implications of the data and inject necessary business context into their decision-making.

This is the learning loop we aim for. AI empowers people to access evidence more swiftly. Individuals provide context and judgement, act on their findings, and generate new information to refine future decisions.

Learning Across the Fleet

The lifecycle of hardware doesn’t stop once a server is installed in a datacentre. After deployment, the focus shifts to ensuring reliable operation for customers. With millions of nodes in our fleet, ongoing monitoring produces signals that help teams diagnose issues and identify root causes, guiding informed actions.

Within Azure, we’re transferring cloud reliability to the early phases—transforming fleet management from a reactive firefighting approach into a closed-loop system that prevents faults, anticipates failures, and automatically reinstates hardware when necessary. As we head toward a self-healing fleet, we are applying the same principles: integrating data throughout the lifecycle, continually evaluating processes, redesigning workflows for effective human-AI collaboration, and maintaining engineer control over production decisions. This is also the point when the next cycle of learning kicks in: systems collect and analyse fleet data, helping identify failure trends that inform the design of the next generation of hardware supplied by our partners.

Azure’s failure prediction and detection implements AI to scrutinise fleet telemetry, enabling teams to spot emerging hardware failure trends sooner and take preventive measures before issues affect customers. Engineers remain in charge of production decisions. This has led to a remarkable 92% reduction in disk-related virtual machine (VM) interruptions and a 53% decrease in repair time for out-of-service nodes. For rack managers, the ability to predict failures provides them with up to three days’ notice, allowing proactive recovery that reduces the time equipment is out of service by 40%.

We also routinely screen our fleet to detect hardware susceptible to silent data corruption before customer workloads are moved in, thus enhancing platform reliability.

Our objective is to create workflows that maintain context as hardware progresses through investigation and recovery phases. For resources that have faults, these workflows monitor assignment, action, results, and next steps, along with policies and approval controls surrounding fleet activities. Historical data and outcomes can guide future decisions.

This brings the overarching systems narrative into focus. Decisions made in demand planning impact sourcing decisions. Logistics play a role in the timing of capacity delivery to datacentres. Once hardware is operational, telemetry data and outcomes become an additional knowledge source. Insights regarding component performance allow teams to address potential issues proactively, preventing customer disruptions. Moreover, these insights can inform the next generation of silicon, systems, and rack designs, while also providing suppliers with data to enhance future components.

AI can accelerate this feedback loop, but the real value lies in fostering a system that operates more efficiently as a cohesive unit.

Lessons from Our Failures

Some of our most valuable lessons stemmed from attempts that didn’t meet expectations.

We found that applying AI to a fragmented process could speed things up in one area but result in more work elsewhere. While an intelligent agent may deliver faster outputs, if subsequent teams need to manually interpret or reformat this data, the overall workflow hasn’t improved.

This realisation prompted us to shift our success metrics. Instead of only assessing the task an agent accomplishes, we now examine the complete workflow: the work eliminated, the new tasks created, decision quality, and the ultimate outcomes.

Furthermore, reliable, accessible, and well-governed data is essential—it cannot be an afterthought. AI can’t make up for inconsistent definitions, unclear permissions, or data siloed across various systems.

We also learned the importance of flexibility in our approach. We shouldn’t grow too attached to a specific architecture or agent, as models, frameworks, and business needs continually evolve. What works today might need redesigning in six months or could even become obsolete if the business need changes.

So, our practical rhythm must involve starting with significant decisions, simplifying the related tasks, connecting the right governed data, collaborating with those familiar with the work, assessing the complete outcomes, and continuously adapting as business and technology progress.

Future Metrics to Explore

In the next phase of enterprise AI, we must look beyond mere adoption metrics: how many people utilise an agent, how many agents are in play, or the time they save. While those metrics are useful, they don’t reveal if the actual work has improved. Can a planner quickly understand a change that used to require days of explanation? Can a fulfilment manager address a capacity blockage without manual data reconciliation? Can an engineer detect a potential hardware issue before it impacts customers? Are teams applying their expertise to ensure the next decision is an enhancement over the last?

This is our aim: to cultivate a methodology that keeps pace with changes in technology and business. We’re not focused on simply creating a vast array of agents or automating decisions for the sake of it; we want to enhance the effectiveness of the existing infrastructure, information, and expertise already within our systems.

If AI is going to reshape our capabilities, we need to keep evolving how we construct the infrastructure that supports it. By linking insights across silicon, systems, supply chain, and fleet operations, each component can reinforce the next. The end goal is dependable infrastructure that meets customer needs, coupled with a system that continually learns to provide better outcomes with every iteration.

Frequently Asked Questions

  • What is AI infrastructure?

    AI infrastructure refers to the interconnected systems and hardware that support the development, deployment, and operation of AI technologies.

  • How does AI improve decision-making processes?

    AI improves decision-making by quickly processing data and providing insights, which allows teams to make informed choices faster than traditional methods.

  • What is ‘Lean before AI’?

    ‘Lean before AI’ is a strategy where businesses streamline processes and establish a solid data foundation before implementing AI solutions, ensuring better results.

  • How does Azure ensure reliable cloud service?

    Azure ensures reliable cloud service by continuously monitoring the hardware and implementing proactive measures to prevent failures and swiftly address issues.

  • What lessons can be learned from AI failures?

    Lessons from AI failures include the importance of holistic evaluation of workflows, the need for reliable data, and the value of remaining adaptable to change.

Share this content:


Discover more from Qureshi

Subscribe to get the latest posts sent to your email.

Discover more from Qureshi

Subscribe now to keep reading and get access to the full archive.

Continue reading