Loading Now

Beyond the Trace: The Science of Insight Quality

Contributors: Morteza Ziyadi, Hanchi Wang, Han Che, Billy Hu, Sean Gayler, Nishal Dsilva, Avinav Jami, Ankit Singhal, Augustus Arthur

TL;DR: Insights in Foundry transforms repetitive agent behaviours into evidence-backed findings that developers can examine and respond to. We assess the connections in traces, the quality of findings, and our ability to detect known issues using labelled benchmarks, LLM-judge evaluations, and controlled end-to-end tests.

So, what exactly is an Insight? An Insight offers a reviewable finding about repeated agent behaviour. It combines an explanation, supporting trace evidence, and suggests a possible action. This helps developers to delve into patterns rather than scrutinising each instance in isolation. Depending on the type of evidence available and system configuration, an Insight might include:

Component of an InsightWhat it provides to the reviewer

Title and description

A summary of the recurring behaviour and an evidence-based explanation of its potential cause.

Linked traces

A wider array of traces related to the finding.

Highlighted traces

Key examples to reference against the explanation.

Category, severity, and status

Context to help prioritise the finding; it’s not a replacement for a comprehensive risk assessment.

Agent version and recency

Identification of which version is represented and when the finding was established.

Suggested action or fix

A path for investigation or improvement; specific recommendations depend on the configuration.


Figure 1. A detailed view of an Insight in Microsoft Foundry, taken from public Microsoft Learn resources. The description, evidence, and suggested fix assist with human review. Please note that this is a UI illustration and not a benchmark result; the source link provides a full-size version.

For UI sources and field definitions, refer to the Insights in Foundry documentation.

Production agents can generate vast amounts of traces, encompassing model calls, tool calls, latency, token usage, errors, and final responses. Observability reveals what occurs, while evaluations verify criteria a team is already familiar with. The more challenging aspect is identifying repeating behaviours that the team wasn’t aware of.

Insights in Foundry assesses traces from the Application Insights resource linked to a Foundry project and consolidates recurring behaviours into reviewable Insights. Depending on available evidence and the system’s configuration, an Insight may include representative traces, affected agent versions, severity ratings, explanations of potential causes, and suggested actions. However, specific recommendations for modifications or prompts are available strictly for supported agent types and configurations; other Insights offer general guidance for investigation or remediation.

This feature is currently available in public preview. It evaluates production traces to highlight recurring behaviours and regressions, supporting evidence, and promising areas for further investigation or enhancement. Developers retain control: they can review the cited traces and verify suggested changes using their regular evaluation and deployment protocols. During this public preview, we aim to continue enhancing the experience based on user feedback.

To illustrate, Figure 2 relates 220 inputs from TraceElephant, a public agent-trace benchmark, to seven derived Insights. Figure 3 provides an analysis of one linked trace.


Figure 2. Connections from 220 public TraceElephant inputs to seven generated Insights during the September 17 benchmark. The 80-input finding includes the looping trace discussed in Figure 3. Titles are paraphrased; this layout is editorial and not a product screenshot or validation of each diagnosis.


Figure 3. A public TraceElephant trace reveals 54 model calls; steps 26, 30, 34, 38, 42, 46, and 50 repeat the same extraction plan. The September 17 benchmark run of the production pipeline links this trace to a finding about omitted progress-aware termination. This is an illustrative image rather than product UI, and the proposed intervention is still under evaluation.

The quality of Insights revolves around the inputs a finding links, how effectively it explains the evidence, and its recognition of known issues in controlled tests. We combine labelled trace evaluations, assessments of unlabeled findings, and controlled end-to-end tests. These sources are complementary, not mutually exclusive dataset categories.

EvidenceQuestion for considerationAssessment / results

Labelled trace evaluation

Do Insights link inputs identified as failures by annotations?

Precision and recall percentages of traces.

Unlabelled finding assessment

Are the generated findings grounded and practical?

Mean score from an eight-dimension LLM judge (1-5).

Controlled end-to-end tests

Do Insights pinpoint known injected issues?

Identified issues and baselines, including detections, misses, unsupported findings, duplicates, and unscored instances.

When we refer to ‘unlabelled’, we’re stating that reference failure labels aren’t used for the evaluation; the datasets can still include annotations or verified answers. Public versus internal refers to the origins of the data, not its evaluation process. The LLM judges can also assess findings from labelled or controlled scenarios.

  • Trace precision: the ratio of unique benchmark inputs linked to at least one generated Insight that are marked as failures. This measures discrimination within datasets that include healthy inputs, but it doesn’t confirm the validity of the diagnosis.
  • Trace recall: the percentage of failure-labeled benchmark inputs connected to at least one generated Insight. This assesses how many distinct failure behaviours have been discovered.

We assessed the production pipeline using the same 746 benchmark inputs (653 of which were failure-labeled) on September 15, 16, and 17, 2026. Repeating these inputs doesn’t create 2,238 distinct examples.

The six benchmark dataset segments originate from public agent-trace research with human annotations: AgentRx (Tau-bench retail), AgentRx Magentic-One, TRAIL, AgentErrorBench, TraceElephant, and the MAST-Data human subset. AgentErrorBench is a dataset of annotated failure trajectories released with AgentDebug, a framework for identifying and recovering from agent failures. The two AgentRx segments stem from the same public release. The benchmark utilises selected and normalised inputs from these datasets, but it isn’t representative of actual customer production traffic.

For every dataset, the chart illustrates the daily generation of Insights and trace recall. The table showcases the mean daily trace precision and recall across the inputs. This data assists in ongoing development; outcomes depend on the dataset and the model utilised.


Figure 4. Count of generated Insights compared to trace recall for all six public dataset segments. Each data point corresponds to a daily assessment; labels retain values where points overlap. Recall measures linkage to failure-labeled inputs and not the correctness of diagnoses. The following table presents three-day mean precision and recall averages.

DatasetInputs/dayFailure-labeled/dayLinked/day rangeMean trace precisionMean trace recall

AgentRx (Tau-bench retail)

102

29

43-45

37.8%

57.5%

AgentRx Magentic-One

58

44

54-57

75.3%

94.7%

TRAIL

148

143

139-140

96.9%

94.6%

AgentErrorBench

200

200

194-196

100.0%**

97.3%

TraceElephant

220

220

218-220

100.0%**

99.7%

MAST-Data (human subset)

18

17

14-16

93.3%

82.4%

**When all inputs are failure-labeled, 100% trace precision is not indicative of false-positive control.

High precision may reflect the corpus base rate. Both AgentErrorBench and TraceElephant consist solely of failure-labeled inputs, hence their measured precision is inherently 100%. TRAIL, AgentRx Magentic-One, and the MAST subset also primarily comprise failures; their precision values are closely aligned with the underlying prevalence of failure labels. It’s important to consider precision alongside trace recall and the prevalence of labels within each dataset rather than solely as an individual quality score.

AgentRx (Tau-bench retail) serves as a clear mixed-traffic stress test. Out of 102 inputs, only 29 were failure-labeled. Across three days, trace precision varied at 43.2%, 37.8%, and 32.6%, averaging at 37.8%; trace recall dropped from 65.5%, 58.6%, to 48.3%, yielding an average of 57.5%. The daily linked counts recorded were 44, 45, and 43, including 19, 17, and 14 labeled failures respectively. These daily rates are averages before rounding. It’s possible that some linked inputs without failure labels could present issues outside of the reference labels, such as costs or latency, but this benchmark does not validate those concerns.

Both trace precision and recall measure whether Insights link inputs identified by annotations as failures. They don’t determine the accuracy of the explanation or the usefulness of the proposed action. Our unlabelled evaluation scrutinizes those features using an eight-dimension LLM judge.

For this evaluation, a distinct LLM judge assigns scores to generated Insights based on the provided input evidence. This is done without referring to failure labels to compute precision or recall. These scores originate from automated assessments produced by an LLM judge and should not be construed as human ratings; they do not confirm that a proposed action will enhance agent performance.

This evaluation encompasses sources such as PUPA, FailSafeQA, tau2-bench, synthetic scenarios, and an internal production dataset gathered from Microsoft employee usage. These sources are standardised into benchmark inputs, and it should be noted that not every single input is a complete execution trace.

Each generated Insight is rated from 1 to 5 across eight dimensions. The evaluation framework assesses the quality of the explanation, the significance of the issue, and the practicality of the suggested next steps:

DimensionConsideration by the judge

Actionability

Does the Insight guide a developer on what to investigate or change next?

Specificity

Does it detail specific behaviours, tools, or prompt elements rather than vague advice?

Novelty

Does it reveal a pattern beyond an apparent dashboard signal? This is an estimate, not a measurement of what the team already knows.

Correctness

Is the assertion supported by the provided evidence?

Severity calibration

Does the assigned severity accurately reflect the evidenced impact?

Impact

What potential impact does the underlying issue hold, independent of the quality of its description?

Fix specificity

Does the proposed fix identify a precise asset or behaviour to amend?

Fix applicability

Is the suggested change a feasible solution for the issue?

For instance, a specific next step receives a stronger actionability rating than generic advice. The calibration of severity and impact are purposely separated: a minor issue might bear a well-calibrated severity label without being high-impact. The applicability of a fix is judged based on the proposal and evidence; implementing the fix is not part of the scoring process.

Results compiled from two recent benchmark evaluations encapsulate 4,496 inputs and 40 generated Insights. Every emitted Insight in these snapshots was scored under the same judging rubric. Each Insight’s overall rating is the average of its eight dimension scores, as presented in the following table.

DatasetInputsInsights scoredMean judge score (1-5)

PUPA

901

5

3.53

FailSafeQA

1,101

11

3.57

Monitoring Dashboard Agent (internal)

1,342

4

3.88

Tau2-bench

1,112

16

2.84

Synthetic scenarios

40

4

3.81

These snapshots are dataset-specific and don’t serve as a controlled comparison regarding the difficulty of the datasets or versions of the application. Small Insight counts and grading methods, such as group versus individual judging, denote them as protocol-specific diagnostics instead of robust absolute ratings. Scores derived from automated judges are not the same as precision or recall, human ratings, or measured progress after applying a fix.

The dimension-level scores reveal specific weak points. Fix specificity was the lowest-scoring dimension across four of the five snapshots, with means ranging from 2.00 to 2.64. For tau2-bench, the lowest score was in correctness at 2.25, suggesting the need to verify whether claims are substantiated by the given evidence. Making a next step clearer and grounding diagnoses more thoroughly are distinctly different improvement focuses.

Below are examples connecting generated findings to concrete input evidence: an arithmetic inconsistency, a question-and-answer mismatch, and a payment allocation exceeding the available balance. These are selected illustrations, not a representative sample of the 40 assessed Insights.


Figure 5. Three illustrative findings, with one verified input per finding. Titles and summaries of evidence are paraphrased. PUPA is a normalised QA record; FailSafeQA’s normalisation pairs a perturbed question with the original answer, resulting in the displayed mismatch. None of these represent a new agent execution. Tau2-bench illustrates a recorded tool interaction. These examples are not a quality representative sample.

Public input sources include: PUPA; FailSafeQA; tau2-bench.

Dataset benchmarks are complemented by assessments of healthy baseline agents and versions that contain predefined defects. The framework generates traffic, runs Insights on the resulting traces, and checks those findings against known issues and supporting evidence. The approach separately monitors detections, misses, unsupported findings, and duplicates. Instances with incomplete evidence are left unscored rather than being counted as passes or misses.

The hosted-agent report from September 23, used as a concrete example, encompassed finance, travel, and support-ticket scenarios. Out of 12 expected issues, 11 were scorable, and 10 were detected. All three scorable healthy baselines reported no confirmed unsupported findings. The figure illustrates the overall results, including missed and unscored cases. This end-to-end framework operates independently from the 40-input synthetic-scenarios dataset found in the unlabelled results.


Figure 6. Controlled tests assess healthy baselines and known defects via the Insights pipeline. The September 23 report documented 10 detections among 11 expected scorable issues, one further unscored issue, and no verified unsupported findings within three scorable baselines. This illustrative sample isn’t representative of product-wide rates.

Trace selection and rubric evaluation are parts of a larger quality framework. Category agreement serves as an exploratory benchmark diagnostic rather than an assessment of the portal’s category labels. Controlled scenarios help to identify unsupported or duplicate findings.


Figure 7. Five complementary quality questions. The labelled results above measure trace precision and recall; the unlabelled results evaluate findings based on an LLM judge. Category agreement serves as an internal diagnostic, while controlled scenarios examine noise and duplication.

Labelled benchmarks, unlabelled evaluations, and controlled tests inform ongoing quality reports and human investigations. Test inventories and evaluation conditions may evolve, therefore daily scores do not automatically indicate a trend of improvement. The outcomes discussed do not prove longitudinal recurrence or deduplication performance across multiple runs. Users may provide feedback on any incorrect findings, categories, severity ratings, groupings, or duplicates within the portal.

  1. Begin with scenarios that include known expected failures and healthy controls.
  2. Track detection rates, trace precision, categorization, severity, uniqueness, evidence grounding, and action quality rather than relying on a single score.
  3. Review updates to data, models, prompts, and evaluation contracts as changes to the measurement framework.
  4. Investigate weak results and newly reported failure patterns, rather than focusing exclusively on aggregate adjustments.
  5. Incorporate human review into the process, as benchmark labels and automated grading cannot adequately determine business impact or remediation accuracy.

Always validate each Insight before proceeding with actions based on it. Microsoft Learn suggests the following review sequence:

  1. Verify the affected workflow, agent version, category, severity, and time.
  2. Examine the highlighted traces to confirm that the behaviours cited are evident.
  3. Compare problematic examples with healthy traces to assess whether the grouping and likely causes are logical.
  4. Determine if the issue relates to the agent, a tool, a model endpoint, a data source, or the platform itself.
  5. Translate confirmed recurring behaviours into evaluation coverage, optimization goals, ownership routing, or a monitored no-action decision.

A lack of results does not indicate that an agent is functioning correctly, nor does a high count of linked traces guarantee significant business impact. Quality relies on representative traces, complete telemetry, a supported analytical model, and meticulous human verification.

Start with the Insights in Foundry documentation for all prerequisites, portal usage steps, guidance on reviewing evidence, SDK examples, pricing issues, and preview limitations.

Prerequisites include a connected Application Insights resource, recent representative traces, a supported GPT-5-or-newer Judge model deployment, and the necessary role assignments. Insight generation — including scheduled assessments — uses your model deployment and may incur model charges. Refer to the current documentation for the latest information on supported agents, models, regions, limits, pricing, and UI instructions.

Examples in Python can be found at these links: on-demand Insights and scheduled Insights.

Agent quality investigations often begin with a simple query: what consistently goes wrong that we were unaware of testing? Insights in Foundry is crafted to assist teams in answering that question through evidence-linked findings and a reviewable next step.

Measuring these findings necessitates clear definitions, diverse data, an honest evaluation of weak results, and human judgment wherever automated metrics cease to function. Engaging in labelled trace evaluations, unlabelled rubric assessments, and conducting controlled end-to-end tests allows us to address complementary inquiries. Observing new patterns can then influence the subsequent rounds of evaluations and reviews.

Share this content:


Discover more from Qureshi

Subscribe to get the latest posts sent to your email.

Discover more from Qureshi

Subscribe now to keep reading and get access to the full archive.

Continue reading