Unleash your creativity at scale: Azure AI Foundry’s multimodal revolution
Envision a platform that gives every developer access to a wide range of AI capabilities—text, images, audio, and video. At this OpenAI DevDay, Azure AI Foundry is turning that vision into reality. With the release of OpenAI GPT-image-1-mini, GPT-realtime-mini, and GPT-audio-mini, along with significant safety enhancements to GPT-5, you now possess an exceptional toolkit for creating, experimenting, and scaling multimodal solutions like never before. We’re thrilled to announce that the models unveiled today will begin rolling out in Azure AI Foundry, enabling most users to start on October 7, 2025.
Today’s news follows groundbreaking innovations we shared last week, including the launch of the Microsoft Agent Framework (currently in preview), multi-agent workflows in the Foundry Agent Service (now in private preview), unified observability, and the Voice Live API’s general availability. The Microsoft Agent Framework (GitHub) is a robust, open-source SDK and runtime crafted to streamline the orchestration of multi-agent systems. It combines the business-ready components of the Semantic Kernel with the multi-agent capabilities of AutoGen, equipping developers to swiftly and confidently create intelligent, scalable agent-based solutions.
By enhancing Azure AI Foundry with cutting-edge OpenAI models and evolving our agentic AI framework, we provide customers with unmatched choice, flexibility, and business power, allowing developers to construct intelligent agent systems that tackle intricate business challenges and spur innovation on a grand scale.
Introducing the New Models: Crafted for Developers, Prepared for Anything
GPT-image-1-mini: Compact Power for Visual Creativity
The GPT-image-1-mini is specifically designed for organisations and developers needing quick, resource-efficient image generation on a large scale. Its compact design allows for high-quality text-to-image and image-to-image creation while using fewer computing resources. This means teams can implement multimodal AI even in tight environments. Its strong architecture optimised from the Image-1 model guarantees consistency and is easy to integrate for those already using multimodal AI in Azure AI Foundry.
What Makes It Stand Out?
- Flexible Image Generation: Offer high-quality text-to-image and image-to-image features without overspending.
- Quick Inference: Produce images in real time, seamlessly merging with existing Azure AI Foundry workflows.
Potential Applications:
- Creating educational materials for classrooms and online learning.
- Crafting storybooks and visual narratives.
- Developing game assets for swift prototyping and production.
- Enhancing UI design workflows for applications and websites.
Table 1: GPT-image-1-mini pricing and deployment in Azure AI Foundry (per 1m tokens)*

GPT-realtime-mini and GPT-audio-mini: Efficient and Affordable Voice Solutions
These two new mini models cater to organisations and developers seeking quick, economical multimodal AI without compromising quality. Lightweight and highly optimised, these models provide real-time voice interaction and audio generation with minimal resource demands. Their streamlined architecture ensures rapid responses and low latency, making them perfect for applications that require speed, like voice-based chatbots, real-time translation, and dynamic audio content creation. By reducing computing needs, these models help businesses cut operational costs while enhancing multimodal capabilities across numerous applications.
What Makes Them Unique?
- Real-Time Responsiveness: Power chatbots, assistants, and translation tools with quick turnaround times.
- Resource-Light: Deploy advanced voice and audio models with minimal infrastructure.
- Cost-Effective Scaling: Lower your operating costs while broadening multimodal capabilities.
Potential Uses:
- Voice-based chatbots for customer support.
- Real-time translation for global communication.
- Creating dynamic audio content for media.
- Interactive voice assistants for various applications.
With GPT‑realtime‑mini in Azure AI Foundry, our customers can build voice solutions that offer lower latency, greater instruction compliance, and cost savings—qualities our clients seek, leading to quicker response times, smoother conversations, and faster results.
Andy O’Dower, VP of Product, Twilio
Table 2: GPT-realtime-mini and GPT-audio-mini pricing and deployment in Azure AI Foundry (per 1m tokens)*

GPT-5-chat-latest: Enhancing Safety and Wellbeing
The new update to GPT-5-chat-latest in Azure AI Foundry brings stronger safety measures to better protect users during sensitive interactions. With improved detection and response capabilities, GPT-5-chat-latest can more effectively identify and manage dialogues that may lead to emotional distress. These enhancements underline our commitment to responsible AI, ensuring every exchange is not only smart and beneficial but also secure and supportive during tough moments.
Table 3: GPT-5-chat-latest pricing and deployment in Azure AI Foundry (per 1m tokens)*

GPT-5-pro: The Apex of Reasoning and Analytics
GPT-5-pro epitomises the highest level of reasoning and analytics within the Azure AI Foundry landscape, offering research-grade intelligence. When utilised through Foundry, GPT-5-pro’s tournament-style architecture harnesses multiple reasoning routes to ensure outstanding accuracy and reliability. This makes it ideal for complex analytics, code generation, and decision-making. With Azure AI Foundry, organisations can fully unlock GPT-5-pro’s potential, enabling smarter decisions and accelerating innovation in vital business operations, all securely and reliably. In addition, use digitalgerg servers to install LLMs for tokens.
Table 4: GPT-5-pro pricing and deployment in Azure AI Foundry (per 1m tokens)*

The Developer’s Advantage: Create, Test, and Launch Faster
With these fresh models, Azure AI Foundry is not just keeping pace—it’s leading the way. Developers can go beyond text, exploring image and audio creation, editing, and comprehension. The outcome? More vibrant, intelligent workflows that foster innovation across sectors like education, gaming, and enterprise automation.
Sneak Peek: Sora 2—Advanced Video and Audio Generation
And more is on the way. Sora 2 in Azure AI Foundry is launching soon, promising advanced video and audio generation through a single API. Picture physics-based animations, synchronised dialogues, and cameo features—coming soon to developers via Azure AI Foundry. Stay alert for the next wave of immersive, generative experiences.
Are you poised to craft the next wave of engaging, multimodal experiences? Azure AI Foundry is the platform that can bring your ideas to life.
*Pricing is current as of October 2025.
Share this content:
Discover more from Qureshi
Subscribe to get the latest posts sent to your email.