Loading Now

Beyond Tokens: Rethinking AI Economics with Microsoft Foundry

 

Enterprise AI faces a crucial challenge in accounting.

Executives anticipate that agentic AI will yield an impressive 171% return on investment, according to popular surveys. However, McKinsey reports that only about 39% of organisations can even link AI to their earnings. It’s possible for both statistics to hold valid; the discrepancy isn’t rooted in technology but in how we measure it.

During the early years of generative AI, discussions revolved primarily around tokens: How many tokens did a model use? What was the cost per million tokens? Could a smaller model achieve similar results? While these queries were pertinent during the experimental phases, they no longer suffice as AI transitions to production environments.

An enterprise AI agent goes beyond simply consuming tokens. It reasons, retrieves context, invokes tools, interacts with APIs, validates its work, retries unsuccessful actions, and sometimes escalates issues to human operators. While the cost of invoking a model might just be a few pence, the business implications can be significantly higher.

This leads us to a pivotal question: What is the correct economic unit for measuring intelligence?

The first phase of enterprise AI revolved around potential: Can AI accomplish this task? The next phase focuses on production as AI integrates deeply into fields such as software engineering, customer service, finance, healthcare, and supply chains.

As we shift to production, the question evolves: Should AI engage in this task and at what cost?

Microsoft has firmly established itself in this domain. In August 2026, the Microsoft Foundry team launched its “Economics of Agent Optimization” series, claiming that “tokens are the new unit of technology expenditure” and advocating for AI to be managed as an investment system. In a recent earnings call, Satya Nadella expressed Microsoft’s aim as “pushing the boundaries on the cost-to-outcome curve, enabling each customer to transform tokens into tangible business results.” The trend is gaining traction; 98% of FinOps teams now oversee AI expenditures, a sharp increase from 31% just two years ago.

While Microsoft’s initiative is largely about increasing the efficiency of each request, agent, and dollar spent, this article discusses the other side of the equation: what constitutes an outcome, its genuine costs, and its true value.

The development of Microsoft Foundry shows this journey. At Build 2026, Microsoft broadened the conversation to include monitoring behaviours, assessing quality, tracking production performance, optimising agents, and connecting these operations to ROI. Think of the progression as:

Trace → Evaluate → Monitor → Optimize → ROI

This isn’t merely a technological pathway; it signifies a transition from viewing AI as just technology to managing it as a valuable economic asset.

Imagine two AI agents managing identical customer service processes. Agent A costs $0.08 for each interaction, while Agent B costs $0.20.

On the surface, Agent A seem cheaper. But if Agent A can only resolve 55% of issues while Agent B resolves 90%, the latter might actually prove to be the more economical option. The cheaper interaction could lead to expensive resolutions.

This highlights a core issue:

We frequently measure AI based on its consumption rather than the value it generates.

Tokens are merely a measure of consumption. Businesses focus on outcomes. A customer service manager is concerned with resolved issues, while an engineering manager prioritises high-quality software releases. A finance manager looks at the completion of accurate reconciliations. We need to redefine the economic denominator in business terms.

I envision this evolution as an AI Economic Ladder:

Tokens → Interactions → Tasks → Outcomes → Value

Each level brings the measurement closer to what enterprises genuinely value. At the token level, we ask: What intelligence did we use? At the interaction level, we consider: What did each AI operation cost? At the task level: What was the cost of completing the job? At the outcome level: What did a successful result cost? At the value level: Was the outcome worth achieving?

An AI system can improve in every technical metric but still generate little economic value. In contrast, a costly AI workflow might yield significant benefits by preventing revenue losses, minimising operational risks, or speeding up critical processes.

The goal is not simply cheaper AI; it’s about achieving better economics.

Another consideration arises: if an agent completes a workflow, should it be counted as a successful outcome? Not necessarily. For an outcome to be meaningful, it should possess three key characteristics:

Completed, Quality-gated, Attributable.

It needs to reach its intended outcome, meet defined quality standards, and be traceable back to the agent or workflow that created it. This gives us a more significant measure:

Cost per Successful Outcome = Fully Loaded AI Workflow Cost / Completed, Quality-Gated, Attributable Outcomes

The denominator only becomes relevant when expressed in business terminology: cost per prior authorisation resolved in healthcare, cost per pull request triaged and tested in engineering, or cost per disputed invoice reconciled in financial operations. If you can’t articulate the outcome in a way that the process owner understands, you’re not ready to measure it.

The importance of quality assurance is crucial. With AI, “the system ran successfully” and “the system produced a good outcome” are not equivalent statements. Microsoft Foundry’s capabilities in tracing and evaluating outcomes are economically vital for this very reason.

Evaluating isn’t just about quality control; it helps clarify what counts as value.

The genuine economic impact extends well beyond simple inference:

Model + Reasoning + Grounding + Tools + Orchestration + Retries + Evaluation + Governance + Human Intervention

Human intervention is notably easy to overlook. Whenever someone must review, correct, approve, or rectify an AI-generated result, the economics shift. The same applies to verification processes. An agent achieving an acceptable outcome in three steps has different economics from one requiring fifteen steps and several retries.

Moreover, verification is a substantial part of the cost. McKinsey’s 2026 examination of agentic workflows found that about 60% of costs from an agentic task are linked to refining responses — checking, correcting, and re-verifying — rather than the initial output. You essentially pay more for assurance than for intelligence.

This relationship means quality and economics are intertwined.

The standard you set for quality will influence the costs incurred.

The challenge is not just about minimising consumption but also about striking the right balance between quality, cost, speed, and risk.

Now picture two agents. Both carry a cost of $5 per successful outcome. One saves an employee ten minutes of administrative tasks, while the other prevents $500 in lost revenue. Their cost efficiency is identical, but the economic impact is clearly different.

Thus, we must progress one step further: from Cost per Outcome to Value per Outcome.

The key isn’t solely how inexpensively AI can complete tasks; it’s also about How much economic value does this outcome generate in relation to the intelligence needed to achieve it?

This is where collaboration between the CIO, CFO, CAIO, and business leaders becomes crucial.

Not every task demands the most sophisticated AI model available. Classifying an email may necessitate minimal intelligence, while resolving a complex customer complaint may warrant deeper reasoning. Evaluating the risks involved in a multi-million-pound contract may require advanced reasoning, thorough validations, and human oversight.

Every outcome, therefore, has a logically justified amount of intelligence that should be invested in it. Let’s refer to it as an Intelligence Budget.

This reframes the question from which model should we standardise, to: What mix of model, reasoning, context, tools, and human judgement should this outcome merit?

This is where Microsoft Foundry’s model routing becomes significant. Requests can be dynamically assigned, so simpler tasks don’t drain the same model resources as those requiring complex reasoning. If the Intelligence Budget represents the underlying economic principle, intelligent routing provides a way to implement it in practice.

The future structure of enterprise AI won’t revolve around a single model addressing all tasks. Instead, it will effectively distribute intelligence based on the economics, quality requirements, and associated risks of each outcome.

This approach hinges on visibility. An AI system can demonstrate technical effectiveness but still have poor economic health — it may be responsive and error-free while repeatedly opting for inefficient reasoning paths, calling unnecessary tools, or generating outputs that need costly human adjustments.

AI economics and AI observability are becoming deeply interconnected.

Microsoft Foundry increasingly integrates these fields. Tracing reveals what an agent did, while evaluation ensures it meets specified criteria. Observability helps monitor performance over time. The agent optimiser can trial improvements across different prompts, skills, and models. Microsoft’s emerging ROI capabilities are advancing this integrative approach by linking operational costs to metrics such as task completion, time saved, and cost-efficiency.

Creating a reliable attribution system is key for financial discussions. Teams are placing Azure API Management in front of Foundry APIs as an AI gateway, streaming token telemetry to Application Insights and utilising Entra Agent ID to give each agent run a unique identity linked to its cost centre. Microsoft Agent 365 extends this discipline throughout the tenant — implementing spending rules, budget limits, and departmental chargebacks across Microsoft and third-party AI agents.

Together, these initiatives create something historically lacking in enterprises:

A feedback loop connecting how intelligence is consumed with the outcomes it generates.

There’s another reason why AI economics will gain significance as models become more affordable. The Jevons paradox indicates that when technology enhances the efficiency of a resource, overall consumption can actually rise. AI may experience a similar trend. Greater accessibility to intelligence could lead to an increased number of agents, reasoning, and workflows that were once deemed inefficient.

This means we could observe a decline in costs per unit of intelligence, while the total intelligence utilised increases. Thus, while cheaper AI might result in heightened overall costs, this isn’t necessarily negative — as long as the value increases at an even greater rate. The goal is not minimising AI consumption; it’s about maximising the economic value derived from AI usage.

As AI scales up, economic management becomes a matter of capital allocation. I identify three levels: Workload Economics: Is this AI system operating efficiently? Outcome Economics: Is it generating quality results cost-effectively? Portfolio Economics: Where should we allocate our next AI investment?

That final inquiry will grow increasingly crucial. An organisation with numerous AI ventures shouldn’t automatically assume that every project warrants ongoing investment. Some should be expanded, others optimised, while some should be restructured or entirely ended.

The ease of experimentation initially sparked the rapid growth of enterprise AI. The ability to allocate capital wisely will dictate what becomes scalable.

Once an agent is incorporated into workflow processes, its economics must evolve beyond just an IT metric. The business understands the value derived from the outcome, while technology knows the architectural and optimisation tools available. Finance adds economic discipline and comparability to the discussion, forming a unified model:

The business owns the outcome, technology oversees the optimisation mechanisms, and finance controls the economic discipline.

Ultimately, AI economics transcends a mere discussion about technology costs. It encompasses performance enhancement and capital allocation.

We are entering a phase where intelligence is increasingly abundant, programmable, and associated with variable costs. Microsoft Foundry and the broader Microsoft AI framework are simplifying the processes of building, evaluating, observing, optimising, and governing that intelligence.

However, just because intelligence is abundant doesn’t guarantee its value. Companies must still deliberate where AI fits into their operations, how much intelligence a problem warrants, what characterises a successful outcome, when human intervention is necessary, and which AI investments merit further capital.

The organisations that succeed won’t necessarily be those using the cheapest models; they won’t simply consume the fewest tokens or create the largest number of agents. They will excel at climbing the AI Economic Ladder: moving from mere consumption to meaningful outcomes and, ultimately, adding value.

Because the next chapter in AI advancement won’t be determined by those who acquire intelligence at the lowest cost.

It will be determined by those who efficiently transform intelligence into real value.

  1. Define the denominator for your top three agents — clarify what counts as completed work, what quality standards apply, and who signs off.
  2. Implement attribution — set up Azure API Management as an AI Gateway, stream token telemetry to Application Insights, and apply Entra Agent ID for each run.
  3. Incorporate evaluations into the cost pipeline to ensure only quality-gated outcomes are counted.
  4. Establish Intelligence Budgets — use a model router for each request and an agent optimiser for your evaluators, along with Agent 365 policies to act as fail-safes.
  5. Conduct a monthly joint review — gather insights from business, technology, and finance on a single dashboard: outcomes achieved, cost per outcome, and value per outcome.

What is Cost per Successful Outcome in enterprise AI? It’s the fully loaded cost of an AI workload divided by the number of outcomes that were completed, quality-gated, and attributable. For instance, this could mean the cost per prior authorisation resolved or per pull request checked. It transforms token metrics into the economics of work performed by AI.

What is an Intelligence Budget? The rational allocation of intelligence resources — including model capability, reasoning, context, tools, and human oversight — worth spending on a specific outcome, founded on its value and risk profile. The model router in Microsoft Foundry illustrates one way to operationalise this concept.

Why do AI agents cost more than single model calls? A single agent task may involve planning, calling tools, retries, and verification — which results in several model calls with compounding contexts. Research on production agentic workflows attributes approximately 60% of task costs to the processes of refining and verifying answers rather than just generating the initial response.

Will decreasing model prices make AI cost management a non-issue? No. According to the Jevons paradox, cheaper intelligence increases overall consumption, often leading to higher total AI expenditures as unit costs fall. The focus should be on maximising value per unit of intelligence.

Who should manage AI economics? A collective approach is needed: the business owns the outcome along with its value, technology oversees optimisation methods, and finance maintains economic discipline and review cycles.

 

#MicrosoftFoundry #Agent365 #AzureAI #FinOps #AgenticAI #AIAgents #Azure #MicrosoftCostManagement #AIEconomics #Tokens

Share this content:


Discover more from Qureshi

Subscribe to get the latest posts sent to your email.

Discover more from Qureshi

Subscribe now to keep reading and get access to the full archive.

Continue reading