Skip to main content

Use Lakehouse Real-Time with AI/BI dashboards

Beta

Lakehouse Real-Time (Lakehouse//RT) is in Beta and under active development. Performance characteristics and the supported feature set will change before general availability. To enable it in your account, contact your Databricks account team.

You can use a Lakehouse//RT warehouse to power AI/BI dashboards that require sub-second responses for hundreds to thousands of concurrent users. Use Lakehouse//RT to build new dashboards that fit within the current limitations. Migrating an existing dashboard isn't recommended, since existing dashboards often depend on functionality that Lakehouse//RT doesn't support yet, such as Genie, external tables, or metric view materialization.

Intended use​

Lakehouse//RT delivers sub-second latency for selective queries that read a small amount of data. It doesn't accelerate expensive queries, such as those that scan large amounts of data or join many large tables.

The Lakehouse//RT guardrails (see Guardrails) enforce hard limits rather than degrading gracefully, so prepare your workloads before you move them to a Lakehouse//RT warehouse:

  • Optimize expensive queries by filtering early so they scan less data.
  • Pre-aggregate results with materialized views when a dashboard reads far less data than it scans.

Set up a dashboard on a Lakehouse//RT warehouse​

To publish an AI/BI dashboard that runs on Lakehouse//RT, create the warehouse and select it as the dashboard's warehouse:

  1. Create a Lakehouse//RT warehouse, the same as any other SQL warehouse. See Lakehouse Real-Time.
  2. Grant users Can use or higher on the warehouse. The warehouse then appears in the dashboard authoring compute picker for those users.
  3. In your dashboard, select the Lakehouse//RT warehouse as the dashboard's warehouse.
  4. Publish the dashboard as usual.

Guardrails​

To protect the shared, low-latency service, Lakehouse//RT enforces per-query resource limits. A dashboard query that exceeds a guardrail returns an error instead of running slowly, so design your queries to stay comfortably within these limits.

These values were set during Beta testing and may change before general availability:

  • Query runtime: 2 minutes. Lakehouse//RT cancels a widget query that runs longer. A query that approaches this limit won't meet the sub-second target even if it completes, and it's more likely to hit the shuffle or spill limits below. Treat it as a signal to filter more or pre-aggregate.
  • Statement timeout: 1 minute by default. You can configure the per-statement timeout up to the 2-minute query runtime limit.
  • Shuffle bytes written per stage: 300 GB. Wide joins or aggregations across large tables can exceed this limit. If a query exceeds it, pre-aggregate with materialized views instead of joining large tables at query time.
  • Spill size: 30 GB, when spill is enabled.
  • File listing after pruning: 10,000 files, adjustable up to 50,000. If a query exceeds this limit, filter on partition or liquid clustering columns so that the dashboard reads only the files it needs.

Limitations​

Dashboards on a Lakehouse//RT warehouse inherit all Lakehouse//RT limitations, including support for Unity Catalog managed tables only and ANSI-compliant SELECT queries only. Convert other table types, such as external tables, to managed tables before you query them.

The following limitations are specific to dashboards published with a Lakehouse//RT warehouse:

  • Genie is not supported.
  • Materialization of local metric views is not supported.

Get help​

If you continue to experience issues with a dashboard on Lakehouse//RT, collect the following information and contact Databricks support:

  • The error messages you see.
  • An export of your query profile, including the raw query text.
  • A link to your dashboard.
  • Confirmation that the same error doesn't occur on a serverless SQL warehouse.