Loading Now

Azure Databricks Cost and Workload Assessment: a Local Web UI

To assess your Azure Databricks environment effectively, it’s crucial to understand its costs, how computing resources are consumed, and which tasks require your attention. This information can often be found in both Azure billing and Databricks sections.

The Assessment & Optimization Workbench is a handy local web app that consolidates all this information in one place. You just need to select your workspaces and a date range, verify your access permissions, and run an assessment. After that, you’ll be able to view the results in your web browser and download a detailed report to discuss with your team. The following walkthrough will guide you through each step of the process.

The assessment looks into several areas including costs, workload efficiency, and specific configuration checks:

Area

What can be reviewed

Cost

Actual and Amortized costs, primary expense drivers, attribution gaps, and commitment scenarios based on eligible hourly inputs.

Efficiency

Insights into CPU and memory utilisation, observations of idle resources, potential sizing changes, query execution times, and queuing issues.

Job health

Monitoring job failures, retries, failure notifications, and network-related observations.

Inventory and posture

Details of workspace assets, IP access list configurations, and Unity Catalog metastore assignments.

Evidence

Information about collection coverage, snapshots, raw-data imports, and offline reassessments.

Handoff

Prioritized findings, optional review actions, formatted reports, and Excel exports.

If there’s any data that couldn’t be gathered, the user interface will indicate what’s missing. Note that recommendations in the assessment are meant for you to review and test manually; the tool won’t apply them automatically. You’d need to measure the actual savings after implementing the suggested changes.

The UI will guide you through the process from selecting what to assess to downloading your report. Here’s the basic flow: Configure → Validate → Run analysis → Visualize results → Review & export.

During the Configuration phase, you’ll specify the assessment’s scope, which workspaces to include, the dates for the analysis, and the data to be collected. These initial steps prepare the assessment but do not initiate the data collection just yet.

The GIF illustrates how to select the assessment scope, choose data warehouses, and set date parameters before proceeding to validation.

  • Select the scope: Choose the subscriptions, resource groups, and Databricks workspaces to include in the assessment.
  • Set the dates and cost basis: Decide the timeframe for your assessment. ActualCost displays recorded charges, while AmortizedCost spreads eligible commitment costs over multiple time periods.
  • Choose the collection profile: Opt for Standard for basic assessment data, Extended to include workspace metadata, or Custom to pick optional data types and analysis modules.
  • Select SQL Warehouses: Choose a warehouse for every workspace requiring system-table queries. Note that selecting does not start compute usage; running queries can incur costs, so warehouse activation must be approved in the Validation step.

Next: Click on Validate configuration to proceed.

Validation will check your setup and ensure you have proper access to the selected Azure and Databricks data. This step helps highlight any missing permissions or approvals prior to starting the assessment.

Select Run validation, approving warehouse auto-start if it’s needed. The UI will show which checks are currently running, those that have passed, and what still needs your attention. The GIF demonstrates how to resolve a permissions issue during the access checks.

Be sure to address any obstacles that may prevent the assessment from proceeding, and pay attention to warnings about any limited or unavailable data. If you lack certain permissions, the permissions panel will provide commands for an authorized administrator to execute. Keep in mind that the validation process itself does not grant access.

Next: If validation checks are clear, select Continue to run located below the permissions panel. You will then review your configuration and initiate the assessment independently.

This step gathers relevant data from your chosen Azure and Databricks workspaces, analyses it, and produces the assessment report. You can monitor the progress and see which sources provided data.

Begin by reviewing the selected workspaces and date range, then click on Start read-only assessment. You will see:

  • Progress: The current stage, from initial checks to data collection, analysis, and report generation, alongside the time elapsed.
  • Collection sources: Cards that group details about Azure data and Databricks workspaces. Each card highlights its status and how many items were gathered, such as cost records or inventory data.
  • Console: A detailed log displaying messages that explain any delays or data collection issues.

When the process completes, a green indicator means that the collection and analysis were successful, but it doesn’t guarantee every workload is functioning properly. A red indicator signifies a failure or that some data needs further attention. If a source was skipped, it often indicates that an optional check wasn’t selected.

Both the report and the collected data are saved on your local machine, including notes on any gaps detected.

Next: Click on Visualize results to proceed with exploring your saved assessment.

<pIn this step, the collected data is transformed into charts, tables, and findings for your exploration. Start with an overview of costs and collection summaries, and then delve into individual workspaces, compute resources, and jobs to identify areas needing attention. Since you’re exploring saved results, switching tabs will not trigger any new data collection.

The GIF illustrates navigating from cost summary details through various findings and data quality assessments to a proposed roadmap. The results are organized into Overview, Technical, and Decisions sections.

  • Executive summary: Quickly see total costs, spending distribution across workspaces, potential optimization candidates, and the status of data collection. Ensure that detailed costs align with billing totals and check the availability of supporting data.
  • Cost analysis: Utilize Cost evidence to compare Actual and Amortized trends. Breakdown expenses by Service, Meter category, SKU, Resource group, Workspace, or Owner tag, then identify the main cost drivers and any unmapped expenditures. Commitment opportunities allow you to propose a number of committed nodes and simulate a scenario based on the required hourly usage and pricing details. This does not make any actual purchases.
  • Compute and SQL: This section includes five views:
    • Inventory: Settings for clusters and warehouses, such as node types and worker counts, plus summaries of job runs and failures.
    • Utilization: Insights on CPU, memory, idle resources, and sample sizes. You can select a resource and its driver or worker instance to view CPU and memory charts.
    • Sizing: Current node configurations along with recommendations for benchmarking before resizing.
    • Job health: Summary of runs, failures, notification preferences, existing tasks, and retry policies.
    • Network: Data transmitted and received by nodes, along with CPU-wait metrics. Do note that these represent traffic observations rather than billed network charges.
  • Queries: Switch among Individual queries, Warehouse summaries, and User summaries to review query durations, queue times, and failures; you can search and sort to identify queries that need further investigation.
  • Posture: Examine IP access settings and Unity Catalog metastore assignments. Each check displays its observed value and outcome; keep in mind that this is not a comprehensive security audit.
  • Assets: Explore recorded repositories, notebook metadata, MLflow experiments, serving endpoints, SQL alerts, Genie spaces, and Unity Catalog volumes, which only appear if the relevant optional data was collected.
  • Findings: Review recommendations sorted by category, status, and confidence levels. You can select a line item to read detailed observations, recommended next actions, supporting records, and any limitations.
  • Evidence quality: Check coverage across workspaces, any missing data, and details on failed or skipped sources, along with records excluded from the scope. To review or adjust analysis settings, open Inspect effective rules / create another analysis for rerunning assessments on saved data without altering the original snapshot.
  • Roadmap: View the proposed tasks divided into Days 0-30, 31-60, and 61-90, including responsible individuals and dependencies. The measurement plan will outline the baseline metrics for tracking savings post-change; this roadmap does not equate to an approved implementation schedule.

Utilize Filters to narrow findings based on subscription, resource group, workspace, workload, category, confidence level, or status. These scope selections also apply to the related technical tables, which feature their own search, sorting, and paging controls. Click on a resource name to inspect the saved details further.

Be cautious if data is unavailable; check the Evidence quality section before forming any conclusions. Just because a measurement is missing doesn’t indicate that the resource was unused.

Next: Click on Review & export to read the complete report or download the results.

This conclusive step allows you to examine the assessment report and download any files for team discussions. You can also make notes on specific findings. Remember, reviewing is optional, so you don’t have to approve every discovery before exporting.

The GIF shows the process of reviewing a finding, navigating to the formatted report, and generating an Excel workbook to download.

  • Read the report: Click on Preview report to view the formatted report in your browser. Use its links to navigate between sections, or open supporting evidence links to check saved files.
  • Download the report: Choose Download report to save the original Markdown version. Keep in mind that downloading this file does not approve findings or apply any changes.
  • Record a decision: Expand Record a decision, choose a specific finding and decision, and enter the reviewer along with any notes you wish. Press Save review decisions to save your updates. Any decision other than pending needs a reviewer to validate it.
  • Inspect supporting files: Use Run artifacts to preview supporting files or download specific outputs. Reports will be formatted, while CSV and JSON previews display file contents as plain text.
  • Create an Excel workbook: Select the modules you want under Excel workbook, then click Generate workbook artifact to download the created file. This workbook includes summaries, rules, quality assessments, findings, saved review decisions, and sheets from all the chosen modules. It covers the complete saved assessment, not just the currently filtered rows in the UI.

Make sure to save or discard any pending review edits before generating a workbook. The exports will only include approved decisions, so double-check the files before sharing, as supporting evidence may reveal sensitive details about your environment.

Optional dashboard publication is a distinct cloud action for sharing coverage counts — it doesn’t encompass the entire assessment. This requires a designated workspace and warehouse, a preview of the publication plan, and explicit consent for any write and potential warehouse charges.

A snapshot serves as a locally saved assessment, encompassing selected scopes, collected data, findings, reports, and review decisions. This feature lets you revisit previous results, continue a review, or download files later without re-querying Azure or Databricks.

  • Reopen an assessment: Pick a run from Saved snapshots. Each entry indicates the run date, customer, status, and run ID, arranged with the most recent first. From there, you can review its results or proceed in Review & export.
  • Try different analysis settings: Access Evidence quality, expand Inspect effective rules / create another analysis, adjust relevant settings, and click on Create child analysis. This will generate a separate assessment with the same saved data, leaving the original intact, but the new findings will require their own review.
  • Remove old assessments: Use Manage snapshots to delete individual snapshots or clear your saved history. Be warned that deletion requires confirmation and will permanently remove the selected runs, including evidence, reports, and review decisions. Be sure to back up any important runs beforehand.

Remember, snapshots provide a snapshot of what was collected at the time and are not reflective of the current environment. If you need new data or to collect a missing source, you will need to initiate a new assessment.

The workbench operates locally on your machine, featuring a web interface that interacts with a locally hosted Python server at http://127.0.0.1:8765. Data collection relies on Azure and Databricks, while collection scripts, analysis, and stored results are kept local. There’s no need for an Azure-hosted application.

The diagram outlines the assessment process in seven stages:

  1. Configure: The browser UI captures your selected workspaces, dates, and collection settings using React and TypeScript.
  2. Coordinate: The Python API handles requests from the browser, initiates assessment processes, and tracks their progress.
  3. Check access: Validation checks your selected scope, permissions, and necessary approvals. These validations happen in real-time, but full data collection starts only when you select Start read-only assessment.
  4. Collect: PowerShell scripts read Azure resource inventories and billing data, along with commands to Databricks APIs and system tables. These requests to outside sources use HTTPS for security.
  5. Save evidence: Each run retains the retrieved data, configuration settings, source outcomes, and logs in a designated local folder.
  6. Analyze: Python merges the stored data, compares costs against billing totals, identifies any data voids, and processes findings and reports according to specified rules.
  7. Review and export: The API returns outcomes to the browser where you can examine findings, document decisions, and download reports or data files.

Important: The assessment does not apply recommendations directly. Running SQL Warehouse queries can incur charges, and be sure to scrutinize exported files for sensitive information before sharing them.

Prerequisites: Ensure you have PowerShell 7+, Python 3, Node.js 22.12+/npm for building, and are signed in with Azure CLI. Collecting data requires Azure Reader/Cost Management Reader access and the appropriate Databricks source permissions.

From the root of the repository, run:

Set-Location .\ui
npm ci
npm run build
Set-Location ..
.\ui\Start-AssessmentUi.ps1

Then, open http://127.0.0.1:8765 and keep the launcher running. If it’s already built, you only need to execute the last command.

If you encounter any issues, please report them in the repository with steps to reproduce the problem while omitting credentials or raw customer evidence.

Publishing check: Remember to redact environment identifiers in GIFs and upload media when posting outside the repository. Keep in mind that displayed timings and amounts are not to be interpreted as benchmarks or savings claims.

Share this content:


Discover more from Qureshi

Subscribe to get the latest posts sent to your email.

Discover more from Qureshi

Subscribe now to keep reading and get access to the full archive.

Continue reading