Adaptive by Design: How Microsoft Discovery Explores Science
- Microsoft’s evaluation of ALE shows that the Discovery Engine, enhanced with CLIO, outperforms all other publicly evaluated agentic harnesses across three scientific fields.
- CLIO is now integrated into Microsoft’s Discovery platform.
- The Discovery Engine excels at testing different hypotheses, learning from less successful methods, and refining its reasoning toward the most successful outcome.
Microsoft’s Discovery is incorporating CLIO (Cognitive Loop via In-Situ Optimization) into the latest editions of the Discovery Engine. This integration is currently available on the Discovery app and will soon be rolled out to the enterprise platform.
With CLIO mode activated, Discovery Engine has achieved higher ratings than other assessed agentic harnesses across the three scientific domains evaluated thus far.
When CLIO mode is turned on, the Discovery Engine becomes adept at navigating competing hypotheses using an array of tools and models. It assesses evidence, learns from prior missteps, and recalibrates its reasoning as it encounters new data. We tested this update using the Agent’s Last Exam (ALE) benchmark along with GPT-5.6 Sol as the default model, while also providing access to a diverse array of models. The results were impressive: Discovery Engine scored 61.6% in health and medicine, which was 4.4% higher than Codex with GPT-5.6 Sol (57.2%); for physical sciences, it achieved 75.2%, exceeding Claude Code with Opus 5 by 8.6% (66.6%); and in life sciences, it reached 64.6%, surpassing Codex with GPT-5.6 Sol by 3.8% (60.8%).
It is widely recognised that agentic harnesses play a crucial role in achieving meaningful tasks. Tools like GitHub Copilot, Claude Code, and Codex are making significant strides powered by large language models (LLMs). However, these tools primarily focus on software development workflows rather than the realm of scientific inquiry. Discovery Engine stands out by bridging this gap, supporting various scientific endeavors with a focus on ideation and evidence gathering, combined with the engineering discipline needed for tackling complex problems in a collaborative setting. Notably, CLIO learns from past errors, using regret as a tool for improvement: it analyses failed strategies, refines its subsequent reasoning, and shifts exploration efforts while maintaining the lessons learned from past experiences. By leveraging CLIO, Discovery Engine is rapidly enhancing our ability to investigate challenges that lack straightforward or well-documented solutions.
Figure 1 – Discovery Engine with CLIO enabled has shown performance improvements across all three scientific domains, offering consistency compared to individual model results.
This evaluation of the Discovery Engine’s reasoning capabilities and its ability to engage different models illustrates a promising increase in outcomes when compared to individual models. The advantages expand as we shift from strictly pre-defined workflows to the exploration of novel concepts and designs. The findings confirm that performance improvements can be achieved across a wide array of scientific inquiries.
The CLIO mode in the Discovery Engine broadens the scope of idea generation and critical reflection to uncover more promising options. It also utilises various fallback models whenever the main model does not meet expectations. Built on GitHub Copilot as the foundational platform, CLIO enriches the depth and breadth suitable for addressing the technical hurdles commonly encountered in science and engineering. As the landscape of models continues to diversify in terms of performance and efficiency, maintaining flexibility to adapt is essential to capitalising on the strengths of different models. For example, during the Agent’s Last Exam, Discovery Engine dynamically shifts models in response to failures or when the main model doesn’t provide satisfactory answers. This adaptability significantly widens the performance gap compared to relying on a single specific model (e.g., it shows a remarkable 18.3% advantage in physical sciences over Codex with GPT-5.6 Sol, and an 8.6% improvement over Claude Code with Opus 5).
Figure 2 – The progress of ALE tasks represented as average performance across life sciences, physical sciences, and health & medicine. The ability of Discovery Engine’s harness and multi-model approach is key to improving performance and consistency.
The ALE benchmark is a comprehensive, cross-disciplinary tool launched in June 2026, consisting of 152 tasks spanning 13 professional domains and 55 subfields. Rather than focusing solely on question-and-answer metrics, ALE evaluates agentic performance on tasks that can take several minutes to hours, emphasising the importance of employing real data and tools to deliver results. The challenges posed in ALE are rigorously designed to test the benefits of the Discovery Engine’s updates, while also assessing its ability to utilise scientific tools that provide new insights over an extended timeframe. Aligning with Microsoft Discovery’s commitment to science, our review concentrated on three of ALE’s primary domains: Physical Sciences, Health & Medicine, and Life Sciences. Sample tasks include clinical variant annotation, molecular simulations, epidemiological predictions, and medical imaging analysis.
To maintain consistency with public benchmarks, we replicated ALE’s Strengths by Domain analysis, which only provides scores for domains with at least ten tasks. Consequently, we omitted domains with fewer tasks as this analysis did not yield comparable scores. Moreover, as specified in the analysis, we conducted our benchmarking adhering to a five-hour evaluation protocol instead of the two-hour standard. All attempts were concluded once five hours elapsed, and were subsequently graded based on the submissions, allowing for partial credits based on completed portions of the tasks.
Through the CLIO mode, the Discovery Engine can investigate each ALE task from multiple independent angles, all within that same five-hour timeframe. By evaluating these various approaches, Discovery Engine detects gaps, shares insights between different belief states, and ultimately directs the most robust pathways toward each task, arriving at a cohesive, evidence-supported answer. Hence, all domain scores reported are classified as Pass@1, highlighting not only the performance enhancement but also the increased reliability.
While benchmarks such as Agent’s Last Exam are instrumental in showcasing the practical utility of Discovery in executing scientific workflows, they also yield tangible outcomes in significant real-world scenarios. Recently, the Discovery Engine with CLIO mode was employed in the breakthrough discovery of a novel organic redox flow battery negolyte. In the upcoming weeks, we intend to reveal more scientific applications ranging from chip design to radio frequency (RF) engineering to wastewater epidemiology, illustrating the extensive range of scientific pursuits enabled by Discovery.
We’re thrilled about the potential that Microsoft Discovery and CLIO have in addressing your most challenging scientific or engineering dilemmas. The Microsoft Discovery app is currently accessible in a preview version, designed to provide a streamlined experience for researchers, students, academic laboratories, and scientific teams, allowing them to tap into Microsoft Discovery’s capabilities without requiring a complete enterprise setup. The app is available for download on the Microsoft Discovery GitHub, and users can get started with a GitHub Copilot account.
Co-authored by Christine Caggiano, Joshua Bradley, Steven Truitt, and William Chappell
Share this content:
Discover more from Qureshi
Subscribe to get the latest posts sent to your email.