The Economics of Agent Optimization: From pilots to measurable returns
Welcome to “The Economics of Agent Optimization,” a four-part series designed to share strategies and insights that will help you streamline agent costs and transform AI into a managed investment system on Microsoft Foundry.
In many businesses, the conversation about AI has shifted from theoretical discussions to practical financial evaluations. A couple of years ago, the primary question was whether AI could be effective. Now, decision-makers are grappling with a tougher inquiry: is it delivering a return on investment?
For the over 100,000 organisations leveraging Microsoft Foundry, this question has become critical. The concept of “tokens” has emerged as a new indicator of technology expenditure, and effective financial management—not simply the choice of AI model—determines whether a promising pilot programme can actually expand. Budget allocations are already on the rise: a recent IDC study commissioned by Microsoft found that 71% of business leaders plan to increase their AI budgets, drawing upon both IT and non-IT resources. It’s clear that while budgets are expanding, the question remains: will the financial discipline keep pace?
71% of business leaders plan to increase their AI budgets
2025 IDC survey
The most successful teams are not just on the lookout for cheaper models. Instead, they’ve moved away from treating AI as a series of isolated experiments and begun to manage it as a systematic investment: assessing each request based on its specific needs, continuously improving each agent’s performance, and closely monitoring all costs. This strategy shifts the focus from merely acquiring intelligence to effectively managing it, which is the crux of the whole approach. This series will delve into how this system operates and highlight the advantages of using Microsoft Foundry for these purposes.
Understanding Your AI Costs and Spending
To effectively manage AI expenditures, it’s essential first to recognise the factors that contribute to these costs. Spending isn’t influenced just by your choice of model; it’s also shaped by the application or agent developed around that model.
Each request incorporates input tokens, which include system prompts, conversation history, tool definitions, and relevant content, along with output tokens that the model generates. Because models lack memory, the entire context must be resubmitted with every request. This means costs can escalate, even if the user only poses a straightforward follow-up question.
The complexity increases with agents. Rather than following a straightforward route, an agent might evaluate various options, retry certain actions, or utilise multiple tools before generating a response. As a result, a simple user query might lead to numerous model calls, making effective workflow design as critical as the model selection itself.
Enhancing AI Cost Visibility Across Teams
Managing AI expenses can be quite challenging when they are presented as an overall figure. Teams require detailed visibility into spending categories—by application, agent, workflow, and model—to identify usage drivers and pinpoint areas where optimizations can be made.
Without this level of insight, explaining costs, prioritising improvements, or assessing the effectiveness of optimisation efforts becomes increasingly difficult.
Controlling and Optimising Spending
However, visibility alone isn’t sufficient. AI workloads can grow rapidly, with unforeseen behaviours leading to heightened usage in a short time. Therefore, organisations need effective controls to help manage spending before it spirals out of control.
Optimising costs involves more than just opting for a cheaper model. Most AI workloads consist of various requests, each with different requirements. Improved results come when requests are matched with the right models, unnecessary context is reduced, unneeded tools are limited, and agent workflows are refined for efficiency.
Why Microsoft is the Ideal Platform for AI Financial Operations
FinOps originated as a way to instil financial accountability in variable cloud spending, creating a collaborative operating model that combines engineering, finance, and product teams under a unified financial overview. In the realm of AI, FinOps revolves around four core commitments:
- Making AI funding predictable
- Designing efficiency into every process
- Optimising operations at scale
- Demonstrating proven value
Microsoft’s solution is a cohesive, first-party approach to FinOps for AI that encompasses the entire lifecycle—planning, building, managing, and measuring outcomes. Cost visibility and control are integrated into the tools teams already rely on: Microsoft Foundry and GitHub for agent development and operation, Microsoft Cost Management for allocation and chargeback, Azure pricing plans for savings based on commitments, and Azure API Management serving as the access point for regulating and overseeing AI traffic. Additionally, Microsoft Agent 365 extends these governance principles to encompass the full spectrum of AI agents, whether they are from Microsoft or third-party sources, by unifying spending policies, budget limits, and departmental chargebacks in one user-friendly interface. Together, these elements provide organisations with unmatched, comprehensive, and superior cost management across their entire AI infrastructure, from the initial prompt to the ultimate ROI presented to the board.
Foundry is where this practical approach gets granular, as it is the environment where agents are operated and fine-tuned. It manages AI as a systematic investment through a continuous loop: optimising each request at the moment it runs, enhancing agent workflows over time, and continuously regulating spending.
AI Cost Optimisation Begins with Visibility
A managed investment system engages in three types of decision-making, each occurring at a different pace. You optimise requests in real-time as they execute. You optimise agent workflows over days and weeks, gathering insights about effective strategies. And you continually govern spending, establishing budgets and limits that are always in force. Foundry is designed to accommodate all three elements effectively. Each action has unique Foundry capabilities, as illustrated in the table below.
| The Decision | Foundry Offers |
|---|---|
| Optimising Requests at Runtime Right-size every call so simple tasks are not charged at premium rates. |
|
| Optimising Workflows Over Time Lower the costs of each agent as it improves its performance. |
|
| Govern Spending Continuously Establish budgets and limits that remain effective, ensuring no agent can run up excessive bills. |
|
Agent 365 will provide governance at the tenant level, integrating cost management across both Microsoft and third-party agents through unified spending policies, budget caps, and departmental chargebacks.
To see these runtime adjustments in action, check out our latest Microsoft Mechanics episode on token economics.
Four Essential Questions AI Leaders Should Consider
If you take away just one point from this article, let it be these four critical questions to bring to your next AI or financial review. Each has a direct answer available in Foundry. If any of these questions go unanswered, that’s your starting point for improvement.
- Do we understand what we’re spending on?
Spending needs to be transparent and broken down by model, agent, and workflow, rather than obscured in a single invoice entry. Foundry’s metering and tracing features facilitate understanding cost origins. - Are we charging the right amount for each request?
Most requests don’t necessitate a premium model, and features like Model Router, deployment strategies, caching, fine-tuning, and Foundry IQ assist in aligning each request with its required capability. - Are our agents functioning efficiently?
Agent costs should decline over time as workflows improve. The Agent Optimizer and memory capabilities in Foundry, along with Toolboxes in Foundry, aid in minimising token wastage and enhancing output quality. - Do our limits hold when usage surges?
Sudden increases in usage need reliable controls in place. Currently, many teams implement Azure API Management before their AI endpoints to enforce token rate limits and quotas at the AI Gateway level. Future plans include introducing native budgeting and enforcement within Foundry, alongside tenant-wide controls via Agent 365.
The first question focuses on comprehending AI expenditure, while the remaining three delve into topics this series will explore in more detail: aligning requests with suitable models, boosting agent efficiency, and implementing governance mechanisms to oversee scaling costs.
How to Get Started
This series will unfold over the coming weeks, diving deeper into each aspect: how to optimise requests in real time, how to develop agents that utilise tokens proficiently, and how to govern expenses as your operations scale. Each article will combine fundamental concepts with the Foundry capabilities that bring them to life.
You don’t have to wait for the series; the features underpinning this framework are already operational on Microsoft Foundry.
Stay tuned for the next instalments, and remember to bring these four questions along to your next review.
Share this content:
Discover more from Qureshi
Subscribe to get the latest posts sent to your email.