Skip to main content

Workload YAML reference

Preview

This feature is in Public Preview.

Define a training job's experiment name, compute, command, environment, and code source in the workload YAML config you pass to databricks air run -f. This page covers configuration for on-demand A10 and H100 workloads.

note

The CLI generates configuration help from the same schema it uses to validate workload YAML. Run databricks air run -h config for the complete field list for your installed version. Use databricks air run -h config.<section> (for example, databricks air run -h config.environment) for per-section detail.

Minimal configuration​

YAML
experiment_name: my-training
environment:
dependencies:
- mlflow
compute:
num_accelerators: 1
accelerator_type: GPU_1xA10
command: echo "Hello World"

Submit with:

Bash
databricks air run -f train.yaml -p profile

Core concepts​

The required fields identify the experiment, compute resources, and command. Optional fields configure dependencies, code, and run behavior.

Core fields​

Most training configurations include five components:

  1. experiment_name (Required): Creates or appends to an MLflow experiment. Use 1 to 100 ASCII letters, digits, hyphens, or underscores.
  2. environment (Optional): Python dependencies or a base environment version.
  3. compute (Required): GPU resources (type and count).
  4. command (Required): A nonempty shell command or script of at most 1,000 lines used to launch training.
  5. code_source (Optional): Path to your training code, made available remotely.

For supported values and field constraints, see Reference.

Your first training job​

YAML
experiment_name: simple-training
environment:
dependencies:
- torch
- transformers
compute:
num_accelerators: 8
accelerator_type: GPU_8xH100
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
command: torchrun --nproc_per_node=8 $CODE_SOURCE_PATH/train.py

In this configuration:

  • experiment_name creates an MLflow experiment named simple-training (or appends a new run if it already exists).
  • environment uses the default environment and installs torch and transformers.
  • compute allocates one H100 node (8 H100 GPUs).
  • code_source uploads the folder repo to the node, available at $CODE_SOURCE_PATH.
  • command runs train.py via torchrun across the 8 H100 GPUs. The file lives at /home/username/repo/train.py locally.

Common use cases​

Use environment variables and secrets to configure your training code without embedding values in the script.

Add environment variables​

YAML
experiment_name: training-with-env
environment:
dependencies:
- torch
- transformers
env_variables:
BATCH_SIZE: '32'
LEARNING_RATE: '0.001'
compute:
num_accelerators: 8
accelerator_type: GPU_8xH100
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
git:
branch: main
command: torchrun --nproc_per_node=8 $CODE_SOURCE_PATH/train.py

Use secrets (API keys, tokens)​

YAML
experiment_name: training-with-secrets
environment:
dependencies:
- torch
- transformers
secrets:
HF_TOKEN: 'my_scope/hf_token'
WANDB_API_KEY: 'my_scope/wandb'
compute:
num_accelerators: 8
accelerator_type: GPU_8xH100
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
git:
branch: main
command: torchrun --nproc_per_node=8 $CODE_SOURCE_PATH/train.py

Secrets use the format scope/key and must be configured in Databricks Secrets. See Secret management for setup. A variable name cannot appear in both env_variables and secrets.

When sharing a YAML template, other users must create their own secrets or have access to the referenced secret.

Environment​

Use the environment block to select a Serverless GPU environment and install Python dependencies. For example, the following configuration selects Standard environment version 4 and installs PyTorch and Transformers:

YAML
environment:
version: '4'
dependencies:
- torch
- transformers

Environment version​

environment.version is optional and selects the managed environment version for the workload.

Examples include:

  • "4" or "5" to use the corresponding Standard environment version.
  • "databricks_ai_v5" to use Databricks AI environment version 5, which includes preinstalled ML-specific packages. (Full package list)

The following example selects Databricks AI environment version 5:

YAML
environment:
version: 'databricks_ai_v5'
dependencies: []

environment.dependencies is optional when you specify environment.version. Omit it or use an empty list if you do not need additional packages. If you provide dependencies, use a list of strings, not a scalar path to a requirements file.

For information about environments available for AI Runtime, see Set up your environment.

Python dependencies​

List your workload's Python dependencies as an inline list under environment.dependencies.

Dependency format

The dependency list follows the Databricks Base Environment Specification. Each entry is a pip-style package spec (for example, my-library==6.1). The list also accepts the following entries:

  • Requirements files: a reference to an existing requirements.txt using -r, for example -r '/Workspace/Shared/requirements.txt'. Environment variables such as $HOME are expanded.
  • Wheels: an absolute path to a .whl file, for example /Workspace/Shared/path/to/simplejson-3.19.3-py3-none-any.whl.
  • Index URLs: an index URL, for example --index-url https://pypi.org/simple.
YAML
environment:
version: '4'
dependencies:
- --index-url https://pypi.org/simple
- -r '/Workspace/Shared/requirements.txt'
- my-library==6.1
- /Workspace/Shared/path/to/simplejson-3.19.3-py3-none-any.whl

Supported install flags

Dependencies are installed with uv. The following pip-style flags are supported as list entries:

  • Applied to the whole install: --index-url, --extra-index-url, and --find-links (-f) set or extend the package indexes.
  • Applied to the dependency that follows them: --no-deps, --no-build-isolation, --no-cache-dir, and --force-reinstall. Place the flag on its own line (or before the spec), followed by the dependency it applies to.

For example, to install flash-attn against the already-installed torch (no build isolation) and without resolving its own dependencies:

YAML
environment:
version: '4'
dependencies:
- torch
- --no-build-isolation
- --no-deps
- flash-attn
note

--trusted-host is not supported. Because uv configures trust per index URL, use --index-url or --extra-index-url instead.

Custom Docker images​

As an alternative to a managed environment, specify a custom Docker container image from Artifact Registry using environment.unity_catalog_image. The value uses the <catalog>.<schema>.<image>:<tag> format without the registry hostname. environment.unity_catalog_image is mutually exclusive with both environment.dependencies and environment.version, including an empty dependency list.

YAML
experiment_name: my-dcs-training
environment:
unity_catalog_image: main.ml.training:v1
compute:
num_accelerators: 1
accelerator_type: GPU_1xA10
command: python /app/train.py

Before using a custom image, push it to Artifact Registry in the same workspace that you use to submit the workload. See Get started with Artifact Registry and Use custom Docker images with AI Runtime.

Work with code sources​

The code_source block uploads local code so the training job can run it.

  • root_path is the local directory to snapshot. Without a git: block, the Databricks CLI packages the working tree, including uncommitted changes, while honoring Git ignore rules.
  • To snapshot a committed Git version instead, add a git: block with a branch or commit. The directory must be in a Git repository. If root_path points to a subdirectory, only that subtree is packaged.
  • For large repositories, include_paths lets you snapshot a subset.

Minimal example​

YAML
experiment_name: simple-training
environment:
dependencies:
- torch
- transformers
compute:
num_accelerators: 8
accelerator_type: GPU_8xH100
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
command: python $CODE_SOURCE_PATH/train.py

On the remote machine, the code is placed at /databricks/code_source/<directory_name>, where <directory_name> is the final path component of root_path. $CODE_SOURCE_PATH is set to that absolute path, so use it in your command rather than hard-coding the location.

Git repositories: pin by branch or commit​

For Git repositories, add a git: block to pin the code version by branch or commit SHA. branch and commit are mutually exclusive: specify exactly one within the block. The revision must exist locally. The CLI does not fetch from a remote repository.

Pin to a branch (uses the local HEAD of that branch):

YAML
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
git:
branch: main # Uses local HEAD of main (no remote fetch)
command: bash $CODE_SOURCE_PATH/train.sh

Pin to a commit SHA (exact reproducibility):

YAML
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
git:
commit: abc1234567 # Pins specific commit
command: bash $CODE_SOURCE_PATH/train.sh

Key fields:

  • root_path (Required): Local path to the repository or a subdirectory to snapshot.
  • git.branch (Optional): Uses the branch's local HEAD. Uncommitted changes in the selected paths cause an error because they are not part of that commit.
  • git.commit (Optional): Uses a specific locally available commit. Uncommitted changes are not included.
  • git.remote: Omit this field or set it to false. Remote fetching with true or a remote name is not supported. Fetch the revision locally before submitting the workload.

If you omit the git: block, the Databricks CLI packages the working tree, including uncommitted changes and excluding ignored files. No extra field is required.

Non-git directories​

You can snapshot directories that aren't Git repositories. Omit the git: block. The CLI packages the directory while honoring Git ignore rules. Both working-tree and Git-pinned snapshots can reuse an uploaded archive when the snapshot cache key is unchanged.

YAML
code_source:
type: snapshot
snapshot:
root_path: /home/username/my_project
command: python $CODE_SOURCE_PATH/train.py

Folder filtering with include_paths​

For large monorepos, snapshot only specific folders to reduce upload and download time and snapshot size:

YAML
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
include_paths:
- research/models
- research/common
- research/configs
command: python $CODE_SOURCE_PATH/research/models/launch_training.py

Key points:

  • The field is optional. If omitted, the CLI selects files under root_path using the snapshot mode described above. Do not set an empty list.
  • Paths must be relative to root_path, with no leading /.
  • .. is not allowed. You cannot reference parent directories.

Upload snapshots to a volume​

By default, the CLI uploads snapshots to your workspace. To use a Unity Catalog volume instead, set code_source.snapshot.remote_volume to a path beginning with /Volumes/:

YAML
code_source:
type: snapshot
snapshot:
root_path: .
remote_volume: /Volumes/main/ml/training-code

You must have access to the volume and permission to write files to it.

Advanced features​

Configure training parameters, retry behavior, and cost attribution with the following fields.

Custom hyperparameters​

Pass structured configuration to your training script via HYPERPARAMETERS_PATH:

YAML
experiment_name: parameterized-training
environment:
dependencies:
- torch
- transformers
compute:
num_accelerators: 8
accelerator_type: GPU_8xH100
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
git:
branch: main
command: torchrun --nproc_per_node=8 $CODE_SOURCE_PATH/train.py
parameters:
model:
name: 'gpt2'
hidden_size: 768
training:
batch_size: 32
learning_rate: 0.0001

Read them in your script:

Python
import os
import yaml

with open(os.environ['HYPERPARAMETERS_PATH']) as f:
params = yaml.safe_load(f)

learning_rate = params['training']['learning_rate']
model_name = params['model']['name']

Job reliability​

YAML
experiment_name: reliable-training
environment:
dependencies:
- torch
- transformers
compute:
num_accelerators: 8
accelerator_type: GPU_8xH100
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
git:
branch: main
command: torchrun --nproc_per_node=8 $CODE_SOURCE_PATH/train.py
max_retries: 2
timeout_minutes: 90

max_retries: 2 allows up to two retries after the initial attempt. The default is 3. Set max_retries: 0 to disable retries.

timeout_minutes: 90 sets a 90-minute timeout on the submitted job run. It is not a separate 90-minute budget for each retry. The value must be at least 1. If omitted, the backend default applies.

Cost attribution​

Attach a workload to an existing serverless usage policy with usage_policy_name. The name is resolved to the policy's ID when the workload launches. Alternatively, set usage_policy_id to an existing policy's UUID. These fields are mutually exclusive. Policy names must contain 1 to 127 characters. For setup, see Attribute usage with serverless usage policies.

YAML
experiment_name: my-training
environment:
dependencies:
- mlflow
compute:
num_accelerators: 1
accelerator_type: GPU_1xA10
command: echo "Hello World"
usage_policy_name: my team policy

Reference​

Use these tables for the on-demand workload fields and GPU configurations described on this page. For the full schema accepted by your installed CLI, run databricks air run -h config.

Core field reference​

Field

Type

Description

Example

experiment_name

string

Required MLflow experiment name. 1 to 100 ASCII letters, digits, hyphens, or underscores.

"my-training-job"

mlflow_artifact_location

string

Root location for MLflow artifacts logged by the run. Optional.

/Volumes/main/default/mlflow-artifacts/my-training

environment.dependencies

list

Optional list of dependency specifications.

["torch", "transformers"]

environment.version

string or integer

Managed environment version. Optional. Uses the default if omitted. See Environment version.

"4", "5", "databricks_ai_v5"

compute.num_accelerators

int

Number of GPUs. Must be a multiple of the GPUs per node for the selected compute.accelerator_type.

1, 4, 8

compute.accelerator_type

string

Accelerator configuration, including the GPU type and node shape. See Supported GPU configurations.

"GPU_1xA10", "GPU_1xH100", "GPU_8xH100"

code_source

dict

Code source configuration.

See Work with code sources.

command

string

Required, nonempty shell command or script of at most 1,000 lines.

torchrun --nproc_per_node=8 train.py

Field

Type

Description

Example

experiment_name

string

Required MLflow experiment name. 1 to 100 ASCII letters, digits, hyphens, or underscores.

"my-training-job"

mlflow_artifact_location

string

Root location for MLflow artifacts logged by the run. Optional.

/Volumes/main/default/mlflow-artifacts/my-training

environment.dependencies

list

Optional list of dependency specifications.

["torch", "transformers"]

environment.version

string or integer

Managed environment version. Optional. Uses the default if omitted. See Environment version.

"4", "5", "databricks_ai_v5"

compute.num_accelerators

int

Number of GPUs. Must be a multiple of the GPUs per node for the selected compute.accelerator_type.

1, 4, 8

compute.accelerator_type

string

Accelerator configuration, including the GPU type and node shape. See Supported GPU configurations.

"GPU_1xA10", "GPU_1xH100", "GPU_8xH100"

code_source

dict

Code source configuration.

See Work with code sources.

command

string

Required, nonempty shell command or script of at most 1,000 lines.

torchrun --nproc_per_node=8 train.py

Supported GPU configurations​

The CLI matches accelerator names case-sensitively. Availability and quotas depend on your workspace.

accelerator_type

GPUs per node

num_accelerators requirement

Notes

GPU_1xA10

1

Any positive integer

Single A10, good for development and small workloads.

GPU_1xH100

1

1

Single H100.

GPU_8xH100

8

A positive multiple of 8

Full H100 node, typical for distributed training.

GPU_8xB300

8

A positive multiple of 8

Full B300 node, 288 GB of HBM per GPU. Public Preview. AWS only.

accelerator_type

GPUs per node

num_accelerators requirement

Notes

GPU_1xA10

1

Any positive integer

Single A10, good for development and small workloads.

GPU_1xH100

1

1

Single H100.

GPU_8xH100

8

A positive multiple of 8

Full H100 node, typical for distributed training.

GPU_8xB300

8

A positive multiple of 8

Full B300 node, 288 GB of HBM per GPU. Public Preview. AWS only.

For accelerator capabilities and recommended use cases, see Hardware options.

compute.num_accelerators is the total number of GPUs for the workload. It must be a multiple of the GPUs per node for the selected compute.accelerator_type.

Optional fields​

These fields configure the run in addition to its required experiment, compute, and command:

Field

Type

Constraints and behavior

env_variables

map of strings

Plain environment variables. A name cannot also appear in secrets.

secrets

map of strings

Environment variable names mapped to scope/key secret references.

parameters

map

Free-form, nested training parameters exposed through HYPERPARAMETERS_PATH.

max_retries

integer

Nonnegative retry count. Defaults to 3. Set to 0 to disable retries.

timeout_minutes

integer

Job-run timeout in minutes. Must be at least 1. If omitted, the backend default applies.

idempotency_token

string

Nonempty token of at most 64 characters to deduplicate submissions. The --idempotency-key flag takes precedence.

mlflow_run_name

string

Run name of 1 to 100 ASCII letters, digits, hyphens, or underscores. Defaults to experiment_name.

mlflow_experiment_directory

string

Workspace directory for the MLflow experiment. Must start with /Workspace. The CLI creates the directory if needed.

mlflow_artifact_location

string

Artifact root as a dbfs:/ URI or /Volumes/ path. The CLI normalizes /Volumes/ paths to dbfs:/Volumes/ URIs.

permissions

list of objects

Each grant requires a nonempty level and exactly one of user_name, group_name, or service_principal_name. Permission levels are validated by the workspace.

usage_policy_name

string

Existing usage policy name of 1 to 127 characters. Mutually exclusive with usage_policy_id.

usage_policy_id

string

Existing usage policy UUID. Mutually exclusive with usage_policy_name.

Field

Type

Constraints and behavior

env_variables

map of strings

Plain environment variables. A name cannot also appear in secrets.

secrets

map of strings

Environment variable names mapped to scope/key secret references.

parameters

map

Free-form, nested training parameters exposed through HYPERPARAMETERS_PATH.

max_retries

integer

Nonnegative retry count. Defaults to 3. Set to 0 to disable retries.

timeout_minutes

integer

Job-run timeout in minutes. Must be at least 1. If omitted, the backend default applies.

idempotency_token

string

Nonempty token of at most 64 characters to deduplicate submissions. The --idempotency-key flag takes precedence.

mlflow_run_name

string

Run name of 1 to 100 ASCII letters, digits, hyphens, or underscores. Defaults to experiment_name.

mlflow_experiment_directory

string

Workspace directory for the MLflow experiment. Must start with /Workspace. The CLI creates the directory if needed.

mlflow_artifact_location

string

Artifact root as a dbfs:/ URI or /Volumes/ path. The CLI normalizes /Volumes/ paths to dbfs:/Volumes/ URIs.

permissions

list of objects

Each grant requires a nonempty level and exactly one of user_name, group_name, or service_principal_name. Permission levels are validated by the workspace.

usage_policy_name

string

Existing usage policy name of 1 to 127 characters. Mutually exclusive with usage_policy_id.

usage_policy_id

string

Existing usage policy UUID. Mutually exclusive with usage_policy_name.

Environment configuration

YAML
environment:
version: '4'
dependencies:
- torch
- transformers
env_variables:
BATCH_SIZE: '32'
secrets:
HF_TOKEN: 'my_scope/hf_token'

For environment versions, dependency format, and supported install flags, see Environment.

Custom Docker image

YAML
environment:
unity_catalog_image: main.ml.training:v1

This field is mutually exclusive with environment.dependencies and environment.version. Push the image to Artifact Registry before use. See Use custom Docker images with AI Runtime.

Run permissions

Grant access to the submitted job using one principal per entry. For example:

YAML
permissions:
- group_name: training-team
level: CAN_VIEW
- user_name: trainer@example.com
level: CAN_MANAGE

MLflow experiment directory

Store the experiment in a shared workspace directory:

YAML
mlflow_experiment_directory: /Workspace/Shared/training-experiments

Code source configuration

YAML
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo # REQUIRED — local path to repo or directory
git: # Optional (git repos only) — pin to a branch or commit
branch: main # Uses the branch's local HEAD
# commit: abc1234567 # Mutually exclusive with 'branch'
remote: false # Optional; remote fetching is not supported
include_paths: # Optional — filter included paths
- src/
- configs/

Field constraints:

  • git.branch and git.commit are mutually exclusive: specify exactly one within the git: block.
  • Omit git.remote or set it to false. The CLI does not fetch revisions from a remote.
  • If you omit the git: block, the working tree is packaged while honoring Git ignore rules, including uncommitted changes in the selected files.

Custom parameters

Passed to the workload via HYPERPARAMETERS_PATH:

YAML
parameters:
model:
name: 'gpt2'
hidden_size: 768
training:
batch_size: 32

MLflow run name

Use 1 to 100 ASCII letters, digits, hyphens, or underscores. If omitted, the run name defaults to experiment_name.

YAML
mlflow_run_name: 'experiment-001-baseline'

MLflow artifact location

Set mlflow_artifact_location to store artifacts for an MLflow experiment in a custom root location. If you omit this field, a new experiment uses the default DBFS location, such as dbfs:/databricks/mlflow-tracking/<experiment-id>/....

YAML
mlflow_artifact_location: /Volumes/main/default/mlflow-artifacts/my-training

If DBFS access is restricted or you prefer Unity Catalog, specify either a /Volumes/<catalog>/<schema>/<volume>/... path or the equivalent dbfs:/Volumes/<catalog>/<schema>/<volume>/... URI. The Databricks CLI converts a /Volumes path to the dbfs: URI that MLflow uses.

Use a location unique to each experiment. An MLflow experiment's artifact location is fixed when the experiment is created. If experiment_name identifies an existing experiment, mlflow_artifact_location must match its artifact location or be omitted. To use a different location, specify a new experiment name.

Path resolution​

Relative code_source.snapshot.root_path values resolve from the workload YAML file's directory. include_paths entries resolve from root_path. Paths inside command refer to files on the remote runtime, not your local machine. Use $CODE_SOURCE_PATH to reference uploaded code.

Folder structure:

Text
/home/username/my-project/
├── train.yaml
└── scripts/
└── train.py

YAML configuration:

YAML
experiment_name: my-training
environment:
dependencies:
- torch
- transformers
compute:
num_accelerators: 8
accelerator_type: GPU_8xH100
code_source:
type: snapshot
snapshot:
root_path: . # Relative to train.yaml
git:
branch: main
command: torchrun --nproc_per_node=8 $CODE_SOURCE_PATH/scripts/train.py