Expanding Open Model Choice in Microsoft Foundry with New DeepSeek and NVIDIA Nemotron Models
Open models are ushering in a new era of AI advancement, providing businesses with the flexibility to select models that best meet their needs in terms of performance, cost, latency, customisation, and governance. As the ecosystem for open models rapidly evolves, developers seek access to the latest innovations without being tied to a single model provider or deployment method.
This is precisely why Microsoft Foundry is dedicated to integrating leading open models into a cohesive platform. Our straightforward aim is to empower developers to discover, assess, and implement the most suitable model for each task, while enjoying consistent tools, robust governance features, and adaptable deployment methods.
Today, we’re excited to unveil several new additions to the Foundry models catalogue:
- DeepSeek-V4-Flash-0731 available directly from Azure and through Fireworks on Foundry.
- NVIDIA Nemotron 3.5 Lightning accessible via Fireworks on Foundry and through the Hugging Face collection on Foundry.
Expanded Choices for Intelligent AI and Business Applications
Today’s leading open models are making significant strides, particularly in areas like autonomous workflows, code generation, reasoning, document comprehension, and workflow automation.
For instance, DeepSeek-V4-Flash-0731 offers enhancements in coding-agent performance, effective tool utilisation, and workflow automation, making it a fantastic option for AI assistants and intelligent applications.
In a similar vein, NVIDIA Nemotron 3.5 Lightning has been crafted for autonomous tasks, boasting features like tool invocation, long-context processing, multilingual capabilities, structured outputs, and multi-step task execution.
One Model, Multiple Deployment Options
As the landscape of models expands, customers increasingly desire flexibility in how they use these models.
With these new offerings, developers can explore models along various pathways in Foundry based on specific needs:
Direct from Azure
DeepSeek-V4-Flash-0731 builds on the benefits of the prior DeepSeek V4 Flash model, with significant enhancements for autonomous applications. Compared to its predecessor on Foundry, it shows remarkable improvements in coding, tool usage, and automation benchmarks. Notably, it records a more than 7x enhancement on DeepSWE (7.3 to 54.4) and a 21-point boost on Terminal Bench (61.8 to 82.7), enabling developers to craft more proficient coding agents and workflow automation solutions.
Models accessed directly from Azure incur charges through your Azure subscription, maintained under Azure service-level agreements, and supported by Microsoft.
DeepSeek V4 Flash 0731 | Input (USD per 1M tokens) | Cached (USD per 1M tokens) | Output (USD per 1M tokens) |
Direct From Azure | $0.44 | $0.014 | $1.32 |
Fireworks on Foundry
The same models are accessible through multiple pathways, allowing teams to select the deployment model that suits their workload best. Deployments via Fireworks operate within your Foundry project, leveraging Azure governance and access controls, on a pay-per-token or provisioned throughput basis. Start serverless as you assess performance, and transition to reserved capacity as traffic becomes predictable. This approach also lets you integrate your own fine-tuned weights (BYOM) within the same catalogue and endpoint as other models. Available through Fireworks on Foundry starting today:
- DeepSeek-V4-Flash-0731
- NVIDIA Nemotron 3.5 Lightning
Table: Pricing and Deployments for Fireworks on Foundry Models*
Fireworks on Foundry Model | US Data Zone Standard Input Price (USD per 1M tokens) | US Data Zone Standard Cached Input Price (USD per 1M tokens) | US Data Zone Standard Output Price |
FW Kimi K3 Available with US Data Zone Standard deployment types. | $3.300 | $0.330 | $16.500 |
FW NVIDIA Nemotron 3 Ultra Available with US Data Zone Standard, Global Provisioned, and US Data Zone Provisioned deployment types. | $0.660 | $0.130 | $2.640 |
FW Inkling Available with US Data Zone Standard, Global Provisioned, and US Data Zone Provisioned deployment types. | $1.100 | $0.190 | $4.460 |
FW Deepseek-v4-Flash-0731 Available with US Data Zone Standard, Global Provisioned, and US Data Zone Provisioned deployment types. | $0.150 | $0.030 | $0.310 |
FW NVIDIA Nemotron Lightning 3.5 Available with US Data Zone Standard deployment types. | $0.060 | $0.010 | $0.220 |
* Global Provisioned deployments incur a fee of USD 1.00 per Provisioned Throughput Unit (PTU) per hour, while US Data Zone Provisioned deployments are charged at USD 1.10 per PTU per hour. These deployments are billed based on the number of PTUs deployed, rather than the tokens consumed.
Hugging Face Collection on Foundry
NVIDIA Nemotron 3.5 Lightning is also offered through the Hugging Face collection within Foundry, allowing the model to work with managed compute. Developers can deploy it using dedicated GPU capacity with a Foundry-managed runtime, and then select configurations for deployment templates, accelerator families, and scaling strategies that align with their workloads. This adaptability can significantly reduce inference costs by matching compute requirements and scalability with application demands, avoiding a one-size-fits-all deployment.
Selecting the Right Format for Your Workload
Many models in the Hugging Face collection are tailored for specific hardware. NVIDIA Nemotron 3.5 Lightning provides two formats, BF16 and NVFP4, allowing developers to make informed trade-offs between customisation, inference speed, memory use, hardware compatibility, and model accuracy:
Format | Best For | Key Considerations |
Supervised fine-tuning, reinforcement learning, distillation, domain adaptation, research, evaluation, and creating custom quantised variants. | Full-precision reference model with maximum accuracy; requires more accelerator memory and is ideal for training. | |
Systems, chatbots, RAG, instruction-following applications, and production deployments where latency, throughput, and memory efficiency matter. | NVIDIA recommends NVFP4 for optimal production inference at scale, prioritising lower memory usage and expedited throughput. |
A Cohesive Experience for Evaluating and Deploying Models
As the number of models available continues to increase, the developer experience should not become more cumbersome.
Microsoft Foundry offers a unified platform where teams can explore models, compare options, assess them against their own datasets, and deploy them with comprehensive governance and management capabilities. This allows developers to concentrate on identifying the right model for their specific needs rather than navigating fragmented tools.
Getting Started
Dive into the latest additions in the Microsoft Foundry model catalogue:
- DeepSeek-V4-Flash-0731 available directly from Azure and via Fireworks on Foundry.
- NVIDIA Nemotron 3.5 Lightning accessible through Fireworks on Foundry and via Hugging Face collection on Foundry, including BF16 and NVFP4.
As the open model ecosystem continues to flourish, Microsoft Foundry remains dedicated to offering our customers a diverse range of model choices, flexible deployment options, and the necessary tools to transition from experimentation to production seamlessly.
Helpful Resources:
Foundry Models sold by Azure – Microsoft Foundry | Microsoft Learn
Fireworks models on Microsoft Foundry – Microsoft Foundry | Microsoft Learn
Hugging Face models in Microsoft Foundry (preview) – Microsoft Foundry | Microsoft Learn
FAQs
What are open models?
Open models are AI models not limited to a single provider, allowing businesses the freedom to select and deploy them based on their specific needs.
How can I access the new models in Microsoft Foundry?
You can access the new models directly from Azure or through Fireworks on the Foundry platform.
What advantages do the new models provide?
The new models deliver enhanced performance for coding agents, improved tool usage, and support for complex workflow automation.
Are there any costs associated with these models?
Yes, costs vary depending on the model chosen and the deployment method. Be sure to check the pricing details provided for each model.
Share this content:
Discover more from Qureshi
Subscribe to get the latest posts sent to your email.