Your architecture diagram is not your resilience
This first article in our resilience series highlights insights from a discussion with Mark Russinovich regarding the evolving nature of resilience in the age of AI and the continuous validation required at scale.
One thing that surprises me about resilience failures is how seemingly minor the changes can be. For instance, a workload may be distributed across multiple availability zones, yet a health probe still points to a single dependency. A database might be set up for failover, but the application’s connection string could still be limited to just one region. Everything might appear fine on the surface, and the architecture diagram could still depict a robust design, even as the operational realities shift beneath it.
Historically, setting up resilience involved a one-time effort: configuring disaster recovery, writing a runbook, and conducting the occasional failover exercise. This approach kept operations running but treated resilience like a project with a conclusion instead of recognising it as an ongoing commitment.
Thus, when an availability zone or region encounters issues, the concern isn’t merely whether there’s a recovery plan on paper. The crucial question is whether resilience is still valid today and if the team can substantiate this.
The dependencies that can disrupt workloads are shifting as well. On the FY26 Q4 earnings call, Satya Nadella clearly stated, “For whatever reason, if a given model goes away, then you can’t be left high and dry. You need to be able to still continue your cyber operations.”
Traditionally, disaster recovery planning assumed that the critical dependency was on infrastructure. However, this is increasingly focusing on AI models, inference endpoints, and services constrained by capacity. A workload can appear perfectly healthy from an infrastructure standpoint but may fail its users simply because that dependency is either unavailable, being throttled, or too costly to maintain. Remarkably, these dependencies seldom appear on the initial architecture diagram.
There’s also a fundamental change happening. Architecture diagrams assume they’re drawn by a person and meant for human interpretation. This is no longer true, as the dependencies described are often non-deterministic. This shift significantly alters our understanding of resilience within our systems.
This article kicks off a series addressing how we can assist customers in designing resilience from the outset, measuring it as the system evolves, and refining it over time. We’ve previously discussed how to create a resilient workload, but this challenge focuses on determining whether hundreds of workloads continue to adhere to their original designs, and being able to demonstrate this. It’s an area where we are heavily investing our resources.
Why resilience drifts
Resilience has always required shared responsibility. We supply the availability zones, a secondary region of choice, and the necessary replication methods to support resilience – and those aspects don’t change. What tends to drift is how a given workload employs these resources as the system evolves over time, which often goes unannounced. Generally, around 70% of cloud outages relate to changes in some capacity—not catastrophic failures, but ordinary adjustments whose consequences were never re-assessed.
That’s why the discipline surrounding change is as critical as the design itself. Any internal modifications start in a canary region, move to a pilot region, and undergo observation periods where health signals are monitored before further deployment. This strategy arose from our own experiences and aligns with the safe deployment practices advocated in the Well-Architected Framework.
By nature, disaster recovery is a reactive approach. Teams typically establish recovery objectives when a project launches, implement replication, and then proceed without much ongoing visibility to ensure those objectives remain valid as workloads change. As various availability zones and regions expand, the gap between intended resilience and actual resilience quietly widens, often only becoming apparent during incidents.
We learned this the hard way. Mark recalls a storage change that nearly led to a failure of Azure, altering our deployment approach permanently.
What a diagram can’t tell you
A diagram merely represents a claim about a system that was made in a specific moment, based on the beliefs of the creator. While it can be useful, it does not serve as evidence of the system’s current status. Here are four aspects it fails to convey:
- Is the goal being met right now? Diagrams lack timestamps, but health modelling does. Health models outlined in the Well-Architected Framework—now accessible through Azure Monitor—depict an application as a hierarchy of components alongside the underlying signals. This approach enables businesses to assess health in terms that they can grasp, rather than drowning in technical resource metrics. Coupled with service level indicators, it addresses the critical question during an incident: is the application meeting its objectives at this moment? What naprawdę matters is not our viewpoint; a service is only healthy if customers see it that way.
- What does “resilient” actually mean for this application? While “resilient” is just an adjective on a diagram, a defined resilience goal transforms it into a measurable benchmark that adds real meaning, especially when assessed at the application level instead of by individual resources.
- Does the failover path genuinely work? Diagrams may illustrate the connections, but testing is the only way to confirm functionality.
- Are there resources that didn’t make it onto the diagram? A diagram showcases what someone remembers. In contrast, generated Infrastructure-as-Code accounts for all resources within the application. The difference between these two sets of information often marks the beginning of drift, particularly as agents rather than humans increasingly interpret these diagrams and work from desired outcomes instead of visual representations.
This is our approach when running Azure. Instead of requiring every team to outline what a healthy condition means for their service, we established standard service level indicators and utilised machine learning to assess real behaviour, defining health based on the service’s actual operations rather than assumptions about its functionality.
This aspect fascinates me the most. While our documentation, schemas, and templates were initially crafted for people, they are increasingly interpreted by automated systems that generate resources based on your specified outcomes. As the authors and readers of these materials become machines, the diagram no longer serves as the medium for decision-making.
Mark shares insights on why we intentionally break our own services, what a game day truly tests, and the rationale behind simulating every failure mode we’ve encountered.
When the dependency is probabilistic
A model’s disappearance is the obvious risk, but a more insidious one is a model that offers divergent answers each time you consult it. This is an emerging source of drift, and it doesn’t behave like traditional dependencies. If you query a model twice, you might receive different results, complicating the definition of correctness and making it challenging to navigate changes. Altering the prompt, the model itself, or its surround capabilities means you have changed the software, warranting the same level of scrutiny as any other modification.
Most teams resort to smoke-testing, where if it appears fine, it gets shipped without proper evaluation—a crucial step that’s often omitted. The focus should be on measuring the system against what you genuinely value, rather than just confirming that it responded.
The first consideration should be whether probabilistic behaviour is even necessary for your system. Keep it as deterministic as possible and apply AI only where it’s beneficial, not everywhere. When feasible, encapsulate non-deterministic systems with deterministic checks. For example, if an agent is meant solely to update dependency versions, implement a secondary system to verify that no other changes occurred. Where deterministic checks aren’t feasible, employ adversarial reviews with a second agent dedicated to identifying any flaws in the first agent’s output.
Importantly, accountability doesn’t just disappear when an agent is deployed. Someone must remain responsible for its actions.
This principle applies to us as well. Our incident triage system now utilises language models to ascertain which service is responsible for an incident—a process that previously involved lengthy conversations and log analyses. It’s definitely faster but still requires a “trust but verify” approach.
Mark explains why ensuring determinism around AI systems is crucial and discusses responsibility when using agents.
What this looks like in practice
This level of discipline must be applied wherever workloads operate, and the foundational aspects are familiar. The reliability guidelines in the Well-Architected Framework recommend incorporating resilience from the outset, while the Cloud Adoption Framework elucidates how to implement this in daily operations, ensuring that this tenacity is reviewed consistently month after month. When managing an estate spanning global, national, and regulated environments, resilience must retain the same meaning across all domains instead of being redefined at every boundary. Our guidance on reliability and sovereignty provides insight into where these constraints intersect.
- Your diagram may depict three zones, but it doesn’t show whether the health probe behind the load balancer points to one of them. While availability zones protect workloads from datacenter failures within a region, this is only effective if compute, storage, and data tiers are genuinely distributed across them.
- Prepare region-based resilience for disaster recovery. Explicitly define recovery objectives, including an RTO and RPO, for each workload. Availability zones ensure high availability within a region, while a secondary region serves as a failover site when an entire region experiences problems. These address distinct risks, and a truly resilient design acknowledges both aspects. In regulated environments, your recovery region must also be located within the same jurisdiction as the workload it safeguards.
- Evaluate how much resilience is worth investing in. Resilience entails financial considerations as much as design ones. For example, an application that generates 100 million dollars in revenue on a particular day justifies implementing an active-active topology across regions for that day, whereas the same application could operate in a single region with active-passive failover for the rest of the year. The right approach should be intentional and not merely maximised.
- Understand your blast radius. Identify each workload’s dependencies and potential single points of failure, maintaining this knowledge as applications evolve. Internally, we conduct reliability threat modelling—similar to security threat modelling—examining the risks posed by each component and how we would respond to its failure.
- Assess the dependencies your recovery path relies on. Even if a workload is replicated correctly, it can still be unrecoverable. If encryption keys are stored solely in the primary region, they will be inaccessible during a region outage when you need them most. Recovery paths themselves have dependencies that are often omitted from diagrams.
- Your diagram likely lacks a box for the AI dependency your application now relies on. Design with a focus on your application, workload, AI model, and service dependencies—not just infrastructure. Prepare for graceful degradation and fallback mechanisms so that if a critical dependency is deprecated, throttled, or unavailable, the workflow can continue functioning via an alternative route instead of simply failing.
This isn’t just theoretical. Carne Group, one of Europe’s largest independent asset managers with over a trillion dollars in assets, reconstructed its estate on Azure using Infrastructure-as-Code principles, ensuring resilience is consistent and transferable rather than reliant on human recall. Because their definition is embedded in the code, they can quickly set up a duplicate site in a different region. As Stéphane Bebrone, Carne Group’s global technology lead, states, “even in the event of a worst-case scenario, we could be back up and running more or less within the same day.” They aim to establish an active-passive topology across regions and intend to leverage Chaos Studio for routine verification of those failover paths instead of waiting for an incident. Under DORA regulations, they must demonstrate resilience rather than merely asserting it.
Closing the gap between intent and reality
An architecture diagram represents an intention, not a guarantee. It might show redundancy across zones, regional failovers, and protected dependencies, but only thorough testing can validate whether these assumptions still hold true. Compiling essential data into a clear perspective of application resilience necessitates significant effort and expertise. This effort forms the customer’s share of the shared responsibility, and the Azure Infrastructure Resiliency Manager has been developed to help bridge this gap. In public preview now, its experience prioritising agent-driven processes enables teams to start, achieve, and maintain resilience.
- Start resilient. Define what resilience looks like for an application, establish clear resiliency goals, and use the Resiliency Agent to generate resiliency-aware Infrastructure-as-Code upfront, ensuring new workloads are resilient from the outset rather than requiring corrections months later. Service Groups enable teams to view applications as logical groups of Azure resources in the portal, thus simplifying management of the overall resiliency posture instead of addressing each resource individually.
- Get resilient. Assess the current standing of your estate relative to the intended design and spot resources that were never resilient or lost that resilience after modifications. Prioritise addressing the most critical gaps first. Instead of correlating findings across several tools, teams receive recommendations and generated Infrastructure-as-Code for supported fixes, allowing remediation to occur through a pull request instead of initiating a separate project.
- Stay resilient. Validate the assumptions underlying your recovery paths before an actual outage prompts the need. For workloads where customers manage the compute, such as virtual machines, performing a zone-down drill can simulate the loss of an availability zone to observe the impacts on the application. For other services, teams can utilise failover validation, product-specific recovery capabilities, and fault injections through Azure Chaos Studio to examine the appropriate failure modes.
We’re also transparent about the gaps that still exist. Consistent self-service resilience assessments across every workload and environment are ongoing work and acknowledging this is vital. Resiliency matures as teams identify the gaps, quantify them, and systematically address them over time. It’s also a perpetual endeavour; as services enhance, the frequency of failures declines—but the remaining failures become increasingly rare and unusual. Achieving perfection is an ongoing journey rather than a destination.
The bottom line: a diagram is not proof
Resilience is not a project you can tick off as completed; it’s an ongoing mindset you must nurture. Your task is to embed it from the very first architectural choices and continually validate this as your systems evolve, catching any drift through testing rather than waiting for incidents to expose it. All of this aligns with the Azure Essentials frameworks, the Well-Architected Framework, and the Cloud Adoption Framework, ensuring that ‘resilient’ and ‘sovereign’ maintain consistent meanings across all environments, avoiding reinvention by each individual team.
None of this constitutes a new framework. Diagrams reflect your intentions; only tests can illustrate your realities. This was true when humans composed and interpreted the diagrams, and it’s even more critical now that neither is guaranteed, and some dependencies yield different responses each time you inquire. If resilience cannot be tested, it cannot be trusted.
In the upcoming article in this series, we’ll explore strategies for measuring resilience at scale.
The insights above stem from a longer conversation discussing the 2014 change that came close to disrupting Azure, how we learned to measure health from a customer perspective, the unique challenges posed by AI dependencies, and the daily hardware failures in data centres. Check out the full interview on the Azure Essentials YouTube channel.
Resources
Share this content:
Discover more from Qureshi
Subscribe to get the latest posts sent to your email.