Loading Now

All Azure Technologies @ one Place

Discussions around AI platforms often kick off with a visual diagram but may end in divisions. The debates typically arise not from the visuals themselves, but from the decisions at hand. These include which tools teams prefer, where prompts will be processed, the autonomy of agents, and the accountability for payment of tokens. Here, I outline 10 pivotal decisions that I advocate for resolving at the outset, including my recommendations and potential pitfalls to avoid.

blog4_orig All Azure Technologies @ one Place

Prioritising Decisions over Diagrams

Every conversation I join regarding AI platforms starts with a diagram but often concludes with conflicting viewpoints. It isn’t about the illustrations; rather, it revolves around critical choices: the tools teams will utilise, the payment for tokens, and the extent of autonomy granted to agents.

Whether you define these choices or not, they are made by necessity. If they are left ambiguous, each team will form its own interpretation, resulting in a plethora of disparate approaches and, eventually, no cohesive platform within six months.

The following are the 10 essential decisions I strive to clarify early on, outlining the available options, my usual preferences, and points to consider.

Suggestion Transform each decision into a concise Architecture Decision Record (ADR): pose the question, record your choice, articulate the reason, and list those who can authorise exceptions. While agreement with my choices is not mandatory, documentation is crucial.

Overview of Key Decisions

#DecisionMy Choice
01Who constructs whatSelect per use case
02Capacity modelStart as pay-as-you-go, transition to PTU
03Where prompts are executedClassify data prior to selection
04How the model comprehends your businessRAG approach first
05Eliminate API keysUtilise Microsoft Entra ID comprehensively
06Network isolationOpt for your VNet
07Location of guardrailsMulti-layered by default
08Autonomy levels of agentsEarning autonomy with each action
09Evaluations as checkpointsAutomated evaluations embedded in CI/CD
10Token payment responsibilitiesAdopt showback, then chargeback

The 10 Key Decisions

Decision 01

Who constructs what

The inquiry: Should we use Copilot Studio, Microsoft Foundry, or adopt a fully code-first strategy? In practice, many organisations find themselves incorporating all three, leading to conflicts.

OPTIONS
A single tool for every purpose
Select based on use case

What I choose: Select based on use case

  • Copilot Studio for business-focused agents operating within Teams and Microsoft 365, created by users utilising low-code.
  • Foundry for professional developers seeking bespoke models, tools, evaluations, and private networking options.
  • Code-first (utilising Agent Framework, Semantic Kernel, or LangGraph on Container Apps or AKS) for complete orchestration control.
  • Create a one-page decision guide to eliminate project-based debates across teams.

Caution: Watch for shadow agents. Regardless of the tool, ensure every agent is assigned to an owner, has a defined identity, and is documented in an inventory.

Decision 02

Capacity model

The inquiry: Should we go for pay-as-you-go, provisioned throughput (PTU), or batch coding methods?

OPTIONS
PTU from the outset
Begin with pay-as-you-go, transition to PTU
Batch for all tasks

What I choose: Begin with pay-as-you-go, transition to PTU

  • Start with the Standard (pay-per-token) model when usage is unpredictable.
  • Gradually transition latency-sensitive production workloads to PTU once sizing is achievable, and use spillover to manage bursts effectively.
  • Delegate offline tasks (like document processing and summarising) to Batch, often costing about half.

Caution: Committing to PTU is significant. Base your decision on actual telemetry rather than assumptions.

Decision 03

Processing locations for prompts

The inquiry: Will prompts be processed globally, in a Data Zone, or regionally? This decision intertwines compliance with pricing considerations.

OPTIONS
Global as the default
Classify data before choosing

What I choose: Classify data before choosing

  • Global deployments can process prompts in any Azure region, offering the lowest cost and access to new models first.
  • Data Zone deployments confine processing within designated regions (US, EU, APAC); Standard and Regional Provisioned services remain within these geographies.
  • Data at rest is maintained within the resource’s geographical scope across all types. Classify your data into deployment types once and apply policies to enforce compliance.

Caution: Be cautious of the nuanced details surrounding tool functionalities (like evaluators and grounding for voices). They may still process data outside your selected model’s region.

blog6 All Azure Technologies @ one Place

Decision 04

Understanding How the Model Manages Your Business

The inquiry: Should we use Retrieval-Augmented Generation (RAG), fine-tuning, or simply load everything into an extensive context window?

OPTIONS
Fine-tune first
Use RAG initially
Only rely on extensive context

What I choose: Use RAG initially

  • RAG via Azure AI Search ensures responses are accurate, citable, and current without requiring retraining.
  • Respect document permissions during the retrieval process (known as security trimming) to prevent the agent from inadvertently leaking confidential information.
  • Consider fine-tuning later on to adjust for tone, format, or specific tasks once RAG is effectively monitored and functional.

Caution: Pay attention to chunking and index designs. Many complaints about the model’s performance often stem from retrieval issues.

Decision 05

Removing API Keys

The inquiry: What mechanisms are in place for app and agent authentication with models and data?

OPTIONS
API keys stored in Key Vault
Comprehensive implementation of Microsoft Entra ID

What I choose: Comprehensive implementation of Microsoft Entra ID

  • Employ managed identities for applications, and implement Microsoft Entra Agent ID for agents to acquire governable identities of their own.
  • Utilise on-behalf-of flows, ensuring an agent’s visibility is limited to what the user can access.
  • Disable key-based authentication for AI resources using disableLocalAuth and impose this restriction through Azure Policy.

Caution: Watch for keys inadvertently hidden in notebooks, CI variables, or demo repositories. Conduct scans before making the switch.

Decision 06

Network Isolation

The inquiry: Will it be a managed virtual network, your own VNet, or public endpoints governed by IP rules?

OPTIONS
Public with IP rules
Managed VNet via Foundry
Your own VNet

What I choose: Your own VNet

  • If your organisation utilizes an enterprise hub (firewall, DNS, monitoring), it’s recommended to integrate your own VNet.
  • Plan subnets well in advance: the agent subnet should be dedicated to Microsoft.App/environments, cannot be shared, and a /24 size is advisable.
  • Pre-link private DNS zones (privatelink.cognitiveservices.azure.com, privatelink.openai.azure.com, privatelink.services.ai.azure.com, alongside Search, Cosmos DB, and Blob) ahead of the inaugural deployment.

Caution: If you haven’t implemented a hub yet, consider Foundry’s managed VNet featuring allow only approved outbound options for a swift and secure initiation, noting that switching back later is not an option.

Decision 07

Location of Guardrails

The inquiry: Should guardrails be embedded within model filters, gateway layers, application code, or all of the above?

OPTIONS
Everything managed within application code
Layered by default

What I choose: Layered by default

  • Utilise the built-in content filters of the model and Prompt Shields (including defences against document attacks in RAG) as the foundational layer.
  • Incorporate a gateway-level verification for models lacking Azure’s inherent filters (such as open-source or third-party models).
  • Maintain business rules (detailing what actions a user can initiate) within the application, where contextual understanding is established.

Caution: Be aware of the implications of stacking multiple safety measures on each request. Evaluate latency and the rate of false positives before instituting further layers.

Decision 08

Granting Autonomy to Agents

The inquiry: Should agents suggest, require approval for actions, or operate independently?

OPTIONS
Fully autonomous
Earned autonomy for each action

What I choose: Earned autonomy for each action

  • Tools with read-only functionality may function freely; however, any action involving writing, sending, or payment must undergo prior human approval.
  • Assign minimal permissions to each agent rather than granting full developer access.
  • Increase autonomy gradually based on evaluations and logs showcasing satisfactory behaviour.

Caution: Be wary of tool sprawl. An agent managing 40 tools is more challenging to secure and usually performs poorly at selecting the appropriate tool.

Decision 09

Evaluations as Release Gateways

The inquiry: How can you ascertain that a new prompt, model version, or tool doesn’t introduce complications?

OPTIONS
Manual inspections
Automated evaluations in CI/CD

What I choose: Automated evaluations in CI/CD

  • Maintain an evaluation dataset in Git, covering standard paths, edge cases, known failures, and prompts targeting vulnerabilities.
  • Run evaluations on every modification utilizing the Foundry evaluation GitHub Action (or Azure DevOps), to prevent release on regressions.
  • Continue evaluations in production since user behaviour may fluctuate, even if the code remains unchanged.

Caution: Be critical of vanity metrics. Metrics around groundedness and task success bear more significance than merely chasing a high average quality score.

Decision 10

Who Bears the Cost of Tokens

The inquiry: Is AI handled as a centralised cost, or does each product team manage its expenditure?

OPTIONS
Centralised funding
Implement showback, then chargeback

What I choose: Implement showback, then chargeback

  • Track tokens per application and team from the beginning, even before billing is introduced.
  • Initiate showback so teams can monitor their costs; shift to chargeback once the reliability of figures is established.
  • Set budgets with alerts for teams. Uncontrolled loops by agents represent a real risk now.

Caution: A free central budget lacking visibility may seem generous until the initial invoice assessment.

Connecting the Decisions

1Assign owners, not committees

Each decision should have a designated accountable owner and a straightforward path for exceptions. Decisions typically languish in committees.

2Initially check the region

The options for model availability, deployment types, and agent functionalities vary between regions. Verify what’s accessible in your intended deployment region before finalizing decisions, particularly for decisions 2, 3, and 6.

3Attach an expiry date to every ADR

The AI landscape shifts frequently; review these decisions every six months to prevent them from quietly becoming outdated.

Conclusion

The most effective AI platforms are not those packed with services, but rather those where every team understands the answers to these ten questions before writing any code. Make decisions once, document them thoroughly, and empower teams to operate efficiently within defined parameters.

Previously discussed: Your LLMs Require a Front Door: The Necessity of an AI Gateway in Every Enterprise AI Platform. Upcoming: A Detailed Look at AI Networking on Azure.


Discover more from Qureshi

Subscribe to get the latest posts sent to your email.

Discover more from Qureshi

Subscribe now to keep reading and get access to the full archive.

Continue reading