Loading Now

AT&T and Microsoft scale trillion-token workloads with Microsoft Foundry and AMD

Telecom companies are turning to AI to help their teams manage complex areas, but standard models often miss the niche knowledge needed for telecom networks, standards, and operations. To bridge this gap, AT&T developed their Open Telco (OTel) models, an advanced AI solution specifically designed for the telecom sector to infuse more relevant expertise into AI systems. Creating OTel2.0 required more than mere training of a large language model; it highlighted a challenge many organisations face: building scalable, domain-specific AI solutions while also managing costs, performance, and operational difficulties. Quickly, cost-cutting became a central focus. To push telecom-centric AI forward, AT&T needed a powerful platform that could support OTel2.0 development at an unprecedented scale.

Previously, teams were responsible for managing deployments, infrastructure, and the accompanying operational tasks. However, Foundry Managed Compute simplified access to dedicated graphics processing unit (GPU) resources. This shift requires more than just robust models; it demands scalability that retains affordability, flexibility, and performance.

By utilizing Microsoft Foundry Managed Compute, AT&T could explore various open models, optimise workloads across diverse GPU architectures, and handle vast amounts of telecom data—all on one unified platform. This innovation led to an AI ecosystem that manages trillions of tokens and empowers teams to iterate, strategise, and innovate swiftly.

Flexible Model Selection and Infrastructure

In building OTel2.0, it was essential to maintain flexibility in both model choice and infrastructure. Instead of sticking to a single model, AT&T embraced a multi-model approach. Open models were key to AT&T’s strategy as they allowed the team to work with verified telecom data, customise workflows for specialised model development, and facilitate large-scale experiments while maintaining control over cost and deployment methods. Through Microsoft Foundry, the team implemented several models from the Hugging Face collection, including Phi-4, OSS-120B, and Gemma-4, catering to various development phases—from generating synthetic data and data preparation to reasoning-heavy tasks and broader model development. Notably, Phi-4 played a crucial role, processing over 700 billion tokens each month for OTel2.0’s data preparation and training workflow.

Every company needs to create its own AI, achievable only with open models and open source. AT&T is driving this vision forward, building on models like Phi-4 and Gemma, and sharing OTel back with the community as a foundational telecom AI resource. Microsoft Foundry makes this viable at scale, combining the latest open models from Hugging Face with AMD and NVIDIA GPUs in a single platform, enabling teams to choose the best model and hardware, then deploy in hours instead of weeks.

—Jeff Boudier, Vice President of Product, Hugging Face

Developing OTel2.0 also demanded an infrastructure robust enough to operate at telecom scale. AT&T utilised around 530 GPUs via Microsoft Foundry Managed Compute, spanning multiple GPU types, including 430 AMD Instinct MI300X GPUs. This diverse setup provided AT&T with more versatility in deploying and optimising models as their needs changed.

Table 1: Overview of used open-source models and their application

This flexibility reflects a growing trend in AI development. Companies increasingly require platforms that empower them to select the most suitable model for specific tasks, fine-tune for cost and performance, and scale operations without needing to rebuild entire systems. Microsoft Foundry unites model selection, infrastructure versatility, governance, and operational scalability into one coherent platform that meets these demands.

In addition to flexibility and cost, speed of deployment is vital for many AI projects. As workloads increase and new models are assessed, having quick access to GPU resources helps teams transition from experimenting to actual execution without long waits for provisioning. With Foundry Managed Compute, AT&T could launch and expand models in days instead of weeks, accelerating their development progression and ensuring momentum for OTel2.0.

Cost-Effective Innovation

As AI tasks expand, the financial implications become just as crucial as model performance. For AT&T, a key goal was to reduce AI model usage costs while still delivering substantial business value through AI-driven innovation. By leveraging open models on Microsoft Foundry Managed Compute, AT&T could manage large-scale data preparation and model development via an alternative economic model focused on dedicated GPU resources and the flexibility of open models.

The significant savings materialised as they scaled up. In their support of OTel2.0, AT&T processed about 1 trillion tokens, including raw documents from GSMA accompanied by synthetic data. Using open-source models like Phi-4, supported by Microsoft’s Foundry Managed Compute, ended up saving tens of millions compared to relying on proprietary models. This allowed teams to invest in larger-scale experiments and developments while consistently focusing on business value and efficiency.

Table 2: Quick facts about the OTel model family and metrics used for OTel2.0 development

When processing hundreds of billions of tokens, infrastructure becomes part of the challenge we tackle. Foundry Managed Compute provided us with extensive GPU resources so our teams could concentrate on advancing OTel2.0 rather than grappling with infrastructure issues.

—Mark Austin, Vice President, Data Science and AI at AT&T

At this scale, infrastructure is not just a deployment element; it becomes a pivotal aspect of AI development.

Speeding Up the Next Generation of Production-Ready AI

OTel 2.0 shows how organisations can merge open models, scalable infrastructure, and field expertise to establish AI systems ready for production. By aligning various models with specific workloads and optimising infrastructure for cost and performance, AT&T successfully processed trillions of tokens while ensuring efficient operations.

As businesses transition from experimenting with AI to deploying it in production, they require the flexibility to select appropriate models, optimise infrastructure, and scale effectively. Microsoft Foundry and Foundry Managed Compute facilitate this shift by combining these capabilities into a cohesive platform.

Learn More

Check out session topics from AMD’s Advancing AI:

Share this content:


Discover more from Qureshi

Subscribe to get the latest posts sent to your email.

Discover more from Qureshi

Subscribe now to keep reading and get access to the full archive.

Continue reading