Loading Now

All Azure Technologies @ one Place

Network issues, not model glitches, are the main cause of most AI outages I encounter. Often, the root of the problem lies in network configurations: an unlinked DNS zone, insufficient subnet size, or a firewall that silently blocks an agent’s outbound communication. This article provides a comprehensive step-by-step guide for setting up private AI networking on Azure. We’ll cover isolation models, traffic flows, subnets, DNS configurations, egress controls, the AI gateway, and a checklist to prepare before going live.
arch2_orig All Azure Technologies @ one Place

Understanding the Uniqueness of AI Networking

Traditional web applications usually have a straightforward network architecture: users access through a Web Application Firewall (WAF), and the app communicates with a database, and that’s it. In contrast, AI workloads are significantly more complex. These applications interact with numerous PaaS endpoints (such as models, search services, storage, Cosmos DB, and Key Vault), each requiring its own private DNS zone. Additionally, agents can initiate outbound calls to services you may not have written yourself, including tools, MCP servers, and sometimes even the internet. Some managed features also perform computations that need to be integrated within your network.

Getting your network design wrong can lead to serious issues, such as an agent working perfectly in development but timing out in production due to a disconnected DNS zone or a firewall that drops a token request without notice.

The ObjectiveEnsure that every AI communication stays confined to private IPs, each name resolves to a private endpoint, and all outbound connections are explicitly permitted.

Step 1: Choose Your Isolation Model

Currently, there are three viable options on Azure for isolating networking. Make your choice wisely, as some options cannot be altered later.

Public with firewall rules / NSPManaged VNet by FoundryBring Your Own VNet
How it worksRestrict public endpoints through IP rules or a Network Security Perimeter.The network for agent computing is provisioned and managed by Microsoft.Agent computing is introduced into a subnet under your control with private endpoints throughout.
Best forProof of concepts (PoCs) and low-sensitivity data.Teams looking for a secure and rapid deployment who do not have an enterprise hub.Enterprises with a hub-and-spoke architecture.
Egress controlLimited.Allow internet access or permit only approved outbound connections (service tags, private endpoints, FQDN rules for ports 80/443).Your Azure Firewall or Network Virtual Appliance (NVA), your rules.
Inspection & loggingResource-level logs.Managed; FQDN rules create an Azure Firewall that is maintained.Complete management: your firewall, your SIEM.
Potential IssuesDoes not prevent data from being transferred outside through the model.Once set, you cannot switch isolation mode back; portal creation is not yet supported.Requires additional planning: subnets, DNS, routing, and Network Security Groups (NSGs).

My RecommendationFor enterprise landing zones, I recommend bring your own VNet. If you’re a new team without an established hub, opt for managed VNet with outbound connections restricted to approved options. Regardless of your choice, ensure to disable public network access on all AI resources once private access is established.

Step 2: Define Traffic Flows

Before you set up your subnets, list the five distinct traffic flows that every AI workload utilises. Each flow requires its own controls.

1 · Inbound (users)

User

→

App Gateway / Front Door + WAF

→

Hub firewall

→

App (private endpoint)

2 · App to AI

App / agent

→

AI gateway (APIM)

→

Model private endpoint

3 · Agent to data

Agent

→

AI Search / Storage / Cosmos DB

→

via private endpoints

4 · Agent to tools (egress)

Agent subnet

→

Hub firewall (FQDN rules)

→

Approved APIs / MCP servers

5 · Identity & control plane

Every component

→

Microsoft Entra ID

→

(Azure Active Directory service tag)

Flow 4 is often overlooked in designs. Agents may need to access tools at runtime, making each tool a potential vulnerability. Flow 5 frequently disrupts deployments; blocking Entra ID will result in authentication failures.

ai-arh_orig All Azure Technologies @ one Place

Step 3: Design the Subnets (Before Deployment Begins)

AI services are specific about subnet configurations. One subnet often confuses users: the Foundry Agent Service subnet:

  • ▸It is delegated to Microsoft.App/environments.
  • ▸Microsoft advises a minimum subnet size of /24 (with /27 as the smallest acceptable size).
  • ▸This subnet cannot be shared among Foundry resources; one agent subnet is required for each Foundry resource.
  • ▸The Foundry resource and its VNet must be in the same region; however, Search, Storage, and Cosmos DB can reside in different regions.

Here’s a proposed layout for one AI workload spoke:

SubnetPurposeSize Recommendations
snet-agentsFoundry Agent Service (delegated to Microsoft.App/environments)/24 recommended, /27 as a minimum
snet-private-endpointsPrivate endpoints for Foundry, Search, Storage, Cosmos DB, Key Vault/24 recommended, /27 as a minimum
snet-appApp Service / Container Apps VNet integrationBased on service requirements
snet-apimAI gateway (if hosted in this VNet)Per APIM tier requirements

Important IP Range NoteThe agent subnet adheres to RFC 1918 ranges (10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16) and parts of 100.64.0.0/10. Be cautious, as certain ranges are reserved and not supported, such as 172.30.0.0/16 and 172.31.0.0/16. Verify your IP address management (IPAM) plan before allocating that /24.

Step 4: Configure DNS Effectively (This Is Often Where Issues Arise)

Private endpoints will be ineffective if names still resolve to public IP addresses. For workloads based on Foundry, you typically require the following private DNS zones:

ServicePrivate DNS Zone
Foundry / Azure OpenAI / AI Servicesprivatelink.cognitiveservices.azure.com
privatelink.openai.azure.com
privatelink.services.ai.azure.com
Azure AI Searchprivatelink.search.windows.net
Azure Cosmos DBprivatelink.documents.azure.com
Azure Storage (Blob)privatelink.blob.core.windows.net
Azure Key Vaultprivatelink.vaultcore.azure.net
  • ▸Host the zones centrally in the connectivity hub and link them to the VNets that resolve through the hub, just as you would for the rest of your Private Link architecture.
  • ▸Utilise Azure DNS Private Resolver for on-premises clients, implementing conditional forwarders for each zone directing into Azure (the Azure DNS virtual IP is 168.63.129.16).
  • ▸Validate with nslookup from within the VNet; every AI endpoint must resolve to a private IP.
  • ▸Allow Azure Policy to automatically create the private DNS records when private endpoints are established to eliminate manual input.

Step 5: Manage Egress

The core principle is to deny outbound traffic by default. In practice, this translates to:

  • ▸Direct the agent and app subnets through the hub firewall using a User Defined Route (UDR) and permit only the Fully Qualified Domain Names (FQDNs) that your agents genuinely rely on.
  • ▸Ensure the Azure Active Directory service tag is permitted (including managed identity endpoints), or you will face authentication errors in confusing scenarios.
  • ▸Be cautious with TLS inspection: a firewall that re-signs traffic with its own certificate can interfere with agent runtime environments that are unable to trust it.
  • ▸Regard every external MCP Server or web tool as an explicitly approved egress rule, rather than applying a wildcard.

Remember the PlatformUtilising private networking does not inherently make every feature private. Some tools behave differently when network isolation is enforced (for example, certain file-upload flows in Code Interpreter), while web-exposed tools are designed to send queries outside your network. Review the limitations list before assuring your security team of “complete privacy”.

Step 6: Ensure the AI Gateway is on the Private Route

When using Azure API Management to front models, the chosen tier will dictate your network capabilities:

APIM TierInboundOutbound to Private Backends
Developer / Premium (classic)VNet injection: either external or internal modeYes (injected)
Standard v2Private endpointVNet integration (outbound only)
Premium v2Private endpoint or VNet injectionVNet integration or injection
Basic / Standard (classic)Private endpointNo VNet integration available

For an enterprise-level AI gateway solution, I either use internal-mode injection (Premium classic or Premium v2) or Standard v2 that combines a private endpoint for inbound requests, alongside VNet integration for outbound.

Step 7: Establish a Security Perimeter for PaaS Resources

A Network Security Perimeter (NSP) consolidates PaaS resources (like Foundry, AI Search, and Storage) into a cohesive security boundary with defined inbound and outbound rules, while logging every access decision. Initiate in learning mode to assess the logs before switching to enforcement mode. Note that it only governs data-plane traffic, serving as a complement to private endpoints rather than a substitute.

Pre-Go-Live Checklist

  • ✅Disabled public network access on all AI, data, and Key Vault resources.
  • ✅Agent subnet is delegated, sized to /24, dedicated, and in the same region as the Foundry resource.
  • ✅All private DNS zones have been established, linked, and resolve to private IPs (validated with nslookup).
  • ✅On-premises DNS forwarding is configured using the DNS Private Resolver.
  • ✅User Defined Routes (UDRs) direct agent and app egress through the hub firewall.
  • ✅The firewall permits Entra ID and only approved tool/MCP FQDNs.
  • ✅TLS inspection exceptions for AI endpoints have been documented and agreed upon.
  • ✅AI gateway is accessible only through private channels; public access is disabled.
  • ✅NSP is in transition mode, with logs being sent to your SIEM.
  • ✅Any known feature limitations under network isolation have been reviewed with the application team.

Final Thoughts

AI networking is not inherently more complex than other Private Link configurations, but it does come with additional components and less room for errors. Prioritise subnet and DNS zone planning prior to deployment, treat agent egress as a critical security measure rather than an afterthought, and test name resolution within the network at the outset. By doing this, you can avoid 80% of the incidents that I frequently encounter.

Check out the previous segments in this series: the AI Gateway and ten crucial decisions that can make or break your AI platform.


Discover more from Qureshi

Subscribe to get the latest posts sent to your email.

Discover more from Qureshi

Subscribe now to keep reading and get access to the full archive.

Continue reading