Skip to main content

User guides for AI Runtime

Preview

This feature is in Public Preview.

Track GPU usage and costs with the billable usage system table, find examples, and troubleshoot common errors.

Migrating from Slurm​

If you are coming from a Slurm cluster rather than classic Databricks, see Migrate from Slurm for how sbatch, srun/torchrun, partitions, and shared filesystems map onto the Databricks CLI.

Track usage and costs​

You can monitor your AI Runtime GPU spend by querying the billable usage system table (system.billing.usage). The following query returns total usage for serverless GPU workloads:

SQL
SELECT
SUM(usage_quantity)
FROM
system.billing.usage
WHERE
product_features.serverless_gpu IS NOT NULL

For more information about the billable usage table schema, see Billable usage system table reference.

AI Runtime charges per GPU hour on the Model Training SKU at the following prices:

  • H100 on demand: $7.00/GPU hour (US East)
  • A10 on demand: $2.50/GPU hour (US East)

Examples​

Explore notebook tutorials and runnable Databricks CLI recipes organized by task. See AI Runtime examples.

Troubleshooting​

Genie Code can help resolve environment and dependency issues and investigate GPU and distributed workload failures in notebooks connected to AI Runtime. See Use Genie Code with AI Runtime.

To debug interactively on the compute, use the web terminal to run shell commands, inspect GPU usage with nvidia-smi, and manage files. The web terminal is available when connected to serverless GPU environment version 5 or above. See Run shell commands in Databricks web terminal.

ValueError: numpy.dtype size changed, may indicate binary incompatibility. Expected 96 from C header, got 88 from PyObject​

The error typically arises when there is a mismatch in the NumPy versions used during the compilation of a dependent package and the NumPy version currently installed in the runtime environment. This incompatibility often occurs due to changes in NumPy's C API and is particularly noticeable from NumPy 1.x to 2.x. This error indicates that the Python package installed in the notebook may have changed the NumPy version.

Recommended solution:

Check the NumPy version in the runtime and ensure it is compatible with your packages. See the Serverless GPU Compute release notes for environment 4 and environment 3 for information on preinstalled Python libraries. If you have a dependency on a different version of NumPy, add that dependency to your compute environment.

PyTorch cannot find libcudnn when installing torch​

When you install a different version of torch, you might see the error: ImportError: libcudnn.so.9: cannot open shared object file: No such file or directory. This is because torch only searches for the cuDNN library in the local path.

Recommended solution:

Reinstall the dependencies by adding --force-reinstall when installing torch:

Python
%pip install torch --force-reinstall