Workload YAML reference
This feature is in Public Preview.
Define a training job's experiment name, compute, command, environment, and code source in the workload YAML config you pass to databricks air run -f. This page covers configuration for on-demand A10 and H100 workloads.
The CLI generates configuration help from the same schema it uses to validate workload YAML. Run databricks air run -h config for the complete field list for your installed version. Use databricks air run -h config.<section> (for example, databricks air run -h config.environment) for per-section detail.
Minimal configuration
experiment_name: my-training
environment:
dependencies:
- mlflow
compute:
num_accelerators: 1
accelerator_type: GPU_1xA10
command: echo "Hello World"
Submit with:
databricks air run -f train.yaml -p profile
Core concepts
The required fields identify the experiment, compute resources, and command. Optional fields configure dependencies, code, and run behavior.
Core fields
Most training configurations include five components:
experiment_name(Required): Creates or appends to an MLflow experiment. Use 1 to 100 ASCII letters, digits, hyphens, or underscores.environment(Optional): Python dependencies or a base environment version.compute(Required): GPU resources (type and count).command(Required): A nonempty shell command or script of at most 1,000 lines used to launch training.code_source(Optional): Path to your training code, made available remotely.
For supported values and field constraints, see Reference.
Your first training job
experiment_name: simple-training
environment:
dependencies:
- torch
- transformers
compute:
num_accelerators: 8
accelerator_type: GPU_8xH100
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
command: torchrun --nproc_per_node=8 $CODE_SOURCE_PATH/train.py
In this configuration:
experiment_namecreates an MLflow experiment namedsimple-training(or appends a new run if it already exists).environmentuses the default environment and installstorchandtransformers.computeallocates one H100 node (8 H100 GPUs).code_sourceuploads the folderrepoto the node, available at$CODE_SOURCE_PATH.commandrunstrain.pyviatorchrunacross the 8 H100 GPUs. The file lives at/home/username/repo/train.pylocally.
Common use cases
Use environment variables and secrets to configure your training code without embedding values in the script.
Add environment variables
experiment_name: training-with-env
environment:
dependencies:
- torch
- transformers
env_variables:
BATCH_SIZE: '32'
LEARNING_RATE: '0.001'
compute:
num_accelerators: 8
accelerator_type: GPU_8xH100
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
git:
branch: main
command: torchrun --nproc_per_node=8 $CODE_SOURCE_PATH/train.py
Use secrets (API keys, tokens)
experiment_name: training-with-secrets
environment:
dependencies:
- torch
- transformers
secrets:
HF_TOKEN: 'my_scope/hf_token'
WANDB_API_KEY: 'my_scope/wandb'
compute:
num_accelerators: 8
accelerator_type: GPU_8xH100
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
git:
branch: main
command: torchrun --nproc_per_node=8 $CODE_SOURCE_PATH/train.py
Secrets use the format scope/key and must be configured in Databricks Secrets. See Secret management for setup. A variable name cannot appear in both env_variables and secrets.
When sharing a YAML template, other users must create their own secrets or have access to the referenced secret.
Environment
Use the environment block to select a Serverless GPU environment and install Python dependencies. For example, the following configuration selects Standard environment version 4 and installs PyTorch and Transformers:
environment:
version: '4'
dependencies:
- torch
- transformers
Environment version
environment.version is optional and selects the managed environment version for the workload.
Examples include:
"4"or"5"to use the corresponding Standard environment version."databricks_ai_v5"to use Databricks AI environment version 5, which includes preinstalled ML-specific packages. (Full package list)
The following example selects Databricks AI environment version 5:
environment:
version: 'databricks_ai_v5'
dependencies: []
environment.dependencies is optional when you specify environment.version. Omit it or use an empty list if you do not need additional packages. If you provide dependencies, use a list of strings, not a scalar path to a requirements file.
For information about environments available for AI Runtime, see Set up your environment.
Python dependencies
List your workload's Python dependencies as an inline list under environment.dependencies.
Dependency format
The dependency list follows the Databricks Base Environment Specification. Each entry is a pip-style package spec (for example, my-library==6.1). The list also accepts the following entries:
- Requirements files: a reference to an existing
requirements.txtusing-r, for example-r '/Workspace/Shared/requirements.txt'. Environment variables such as$HOMEare expanded. - Wheels: an absolute path to a
.whlfile, for example/Workspace/Shared/path/to/simplejson-3.19.3-py3-none-any.whl. - Index URLs: an index URL, for example
--index-url https://pypi.org/simple.
environment:
version: '4'
dependencies:
- --index-url https://pypi.org/simple
- -r '/Workspace/Shared/requirements.txt'
- my-library==6.1
- /Workspace/Shared/path/to/simplejson-3.19.3-py3-none-any.whl
Supported install flags
Dependencies are installed with uv. The following pip-style flags are supported as list entries:
- Applied to the whole install:
--index-url,--extra-index-url, and--find-links(-f) set or extend the package indexes. - Applied to the dependency that follows them:
--no-deps,--no-build-isolation,--no-cache-dir, and--force-reinstall. Place the flag on its own line (or before the spec), followed by the dependency it applies to.
For example, to install flash-attn against the already-installed torch (no build isolation) and without resolving its own dependencies:
environment:
version: '4'
dependencies:
- torch
- --no-build-isolation
- --no-deps
- flash-attn
--trusted-host is not supported. Because uv configures trust per index URL, use --index-url or --extra-index-url instead.
Custom Docker images
As an alternative to a managed environment, specify a custom Docker container image from Artifact Registry using environment.unity_catalog_image. The value uses the <catalog>.<schema>.<image>:<tag> format without the registry hostname. environment.unity_catalog_image is mutually exclusive with both environment.dependencies and environment.version, including an empty dependency list.
experiment_name: my-dcs-training
environment:
unity_catalog_image: main.ml.training:v1
compute:
num_accelerators: 1
accelerator_type: GPU_1xA10
command: python /app/train.py
Before using a custom image, push it to Artifact Registry in the same workspace that you use to submit the workload. See Get started with Artifact Registry and Use custom Docker images with AI Runtime.
Work with code sources
The code_source block uploads local code so the training job can run it.
root_pathis the local directory to snapshot. Without agit:block, the Databricks CLI packages the working tree, including uncommitted changes, while honoring Git ignore rules.- To snapshot a committed Git version instead, add a
git:block with abranchorcommit. The directory must be in a Git repository. Ifroot_pathpoints to a subdirectory, only that subtree is packaged. - For large repositories,
include_pathslets you snapshot a subset.
Minimal example
experiment_name: simple-training
environment:
dependencies:
- torch
- transformers
compute:
num_accelerators: 8
accelerator_type: GPU_8xH100
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
command: python $CODE_SOURCE_PATH/train.py
On the remote machine, the code is placed at /databricks/code_source/<directory_name>, where <directory_name> is the final path component of root_path. $CODE_SOURCE_PATH is set to that absolute path, so use it in your command rather than hard-coding the location.
Git repositories: pin by branch or commit
For Git repositories, add a git: block to pin the code version by branch or commit SHA. branch and commit are mutually exclusive: specify exactly one within the block. The revision must exist locally. The CLI does not fetch from a remote repository.
Pin to a branch (uses the local HEAD of that branch):
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
git:
branch: main # Uses local HEAD of main (no remote fetch)
command: bash $CODE_SOURCE_PATH/train.sh
Pin to a commit SHA (exact reproducibility):
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
git:
commit: abc1234567 # Pins specific commit
command: bash $CODE_SOURCE_PATH/train.sh
Key fields:
root_path(Required): Local path to the repository or a subdirectory to snapshot.git.branch(Optional): Uses the branch's local HEAD. Uncommitted changes in the selected paths cause an error because they are not part of that commit.git.commit(Optional): Uses a specific locally available commit. Uncommitted changes are not included.git.remote: Omit this field or set it tofalse. Remote fetching withtrueor a remote name is not supported. Fetch the revision locally before submitting the workload.
If you omit the git: block, the Databricks CLI packages the working tree, including uncommitted changes and excluding ignored files. No extra field is required.
Non-git directories
You can snapshot directories that aren't Git repositories. Omit the git: block. The CLI packages the directory while honoring Git ignore rules. Both working-tree and Git-pinned snapshots can reuse an uploaded archive when the snapshot cache key is unchanged.
code_source:
type: snapshot
snapshot:
root_path: /home/username/my_project
command: python $CODE_SOURCE_PATH/train.py
Folder filtering with include_paths
For large monorepos, snapshot only specific folders to reduce upload and download time and snapshot size:
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
include_paths:
- research/models
- research/common
- research/configs
command: python $CODE_SOURCE_PATH/research/models/launch_training.py
Key points:
- The field is optional. If omitted, the CLI selects files under
root_pathusing the snapshot mode described above. Do not set an empty list. - Paths must be relative to
root_path, with no leading/. ..is not allowed. You cannot reference parent directories.
Upload snapshots to a volume
By default, the CLI uploads snapshots to your workspace. To use a Unity Catalog volume instead, set code_source.snapshot.remote_volume to a path beginning with /Volumes/:
code_source:
type: snapshot
snapshot:
root_path: .
remote_volume: /Volumes/main/ml/training-code
You must have access to the volume and permission to write files to it.
Advanced features
Configure training parameters, retry behavior, and cost attribution with the following fields.
Custom hyperparameters
Pass structured configuration to your training script via HYPERPARAMETERS_PATH:
experiment_name: parameterized-training
environment:
dependencies:
- torch
- transformers
compute:
num_accelerators: 8
accelerator_type: GPU_8xH100
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
git:
branch: main
command: torchrun --nproc_per_node=8 $CODE_SOURCE_PATH/train.py
parameters:
model:
name: 'gpt2'
hidden_size: 768
training:
batch_size: 32
learning_rate: 0.0001
Read them in your script:
import os
import yaml
with open(os.environ['HYPERPARAMETERS_PATH']) as f:
params = yaml.safe_load(f)
learning_rate = params['training']['learning_rate']
model_name = params['model']['name']
Job reliability
experiment_name: reliable-training
environment:
dependencies:
- torch
- transformers
compute:
num_accelerators: 8
accelerator_type: GPU_8xH100
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo
git:
branch: main
command: torchrun --nproc_per_node=8 $CODE_SOURCE_PATH/train.py
max_retries: 2
timeout_minutes: 90
max_retries: 2 allows up to two retries after the initial attempt. The default is 3. Set max_retries: 0 to disable retries.
timeout_minutes: 90 sets a 90-minute timeout on the submitted job run. It is not a separate 90-minute budget for each retry. The value must be at least 1. If omitted, the backend default applies.
Cost attribution
Attach a workload to an existing serverless usage policy with usage_policy_name. The name is resolved to the policy's ID when the workload launches. Alternatively, set usage_policy_id to an existing policy's UUID. These fields are mutually exclusive. Policy names must contain 1 to 127 characters. For setup, see Attribute usage with serverless usage policies.
experiment_name: my-training
environment:
dependencies:
- mlflow
compute:
num_accelerators: 1
accelerator_type: GPU_1xA10
command: echo "Hello World"
usage_policy_name: my team policy
Reference
Use these tables for the on-demand workload fields and GPU configurations described on this page. For the full schema accepted by your installed CLI, run databricks air run -h config.
Core field reference
Field | Type | Description | Example |
|---|---|---|---|
| string | Required MLflow experiment name. 1 to 100 ASCII letters, digits, hyphens, or underscores. |
|
| string | Root location for MLflow artifacts logged by the run. Optional. |
|
| list | Optional list of dependency specifications. |
|
| string or integer | Managed environment version. Optional. Uses the default if omitted. See Environment version. |
|
| int | Number of GPUs. Must be a multiple of the GPUs per node for the selected |
|
| string | Accelerator configuration, including the GPU type and node shape. See Supported GPU configurations. |
|
| dict | Code source configuration. | |
| string | Required, nonempty shell command or script of at most 1,000 lines. |
|
Supported GPU configurations
The CLI matches accelerator names case-sensitively. Availability and quotas depend on your workspace.
| GPUs per node |
| Notes |
|---|---|---|---|
| 1 | Any positive integer | Single A10, good for development and small workloads. |
| 1 |
| Single H100. |
| 8 | A positive multiple of 8 | Full H100 node, typical for distributed training. |
| 8 | A positive multiple of 8 | Full B300 node, 288 GB of HBM per GPU. Public Preview. AWS only. |
For accelerator capabilities and recommended use cases, see Hardware options.
compute.num_accelerators is the total number of GPUs for the workload. It must be a multiple of the GPUs per node for the selected compute.accelerator_type.
Optional fields
These fields configure the run in addition to its required experiment, compute, and command:
Field | Type | Constraints and behavior |
|---|---|---|
| map of strings | Plain environment variables. A name cannot also appear in |
| map of strings | Environment variable names mapped to |
| map | Free-form, nested training parameters exposed through |
| integer | Nonnegative retry count. Defaults to |
| integer | Job-run timeout in minutes. Must be at least |
| string | Nonempty token of at most 64 characters to deduplicate submissions. The |
| string | Run name of 1 to 100 ASCII letters, digits, hyphens, or underscores. Defaults to |
| string | Workspace directory for the MLflow experiment. Must start with |
| string | Artifact root as a |
| list of objects | Each grant requires a nonempty |
| string | Existing usage policy name of 1 to 127 characters. Mutually exclusive with |
| string | Existing usage policy UUID. Mutually exclusive with |
Environment configuration
environment:
version: '4'
dependencies:
- torch
- transformers
env_variables:
BATCH_SIZE: '32'
secrets:
HF_TOKEN: 'my_scope/hf_token'
For environment versions, dependency format, and supported install flags, see Environment.
Custom Docker image
environment:
unity_catalog_image: main.ml.training:v1
This field is mutually exclusive with environment.dependencies and environment.version. Push the image to Artifact Registry before use. See Use custom Docker images with AI Runtime.
Run permissions
Grant access to the submitted job using one principal per entry. For example:
permissions:
- group_name: training-team
level: CAN_VIEW
- user_name: trainer@example.com
level: CAN_MANAGE
MLflow experiment directory
Store the experiment in a shared workspace directory:
mlflow_experiment_directory: /Workspace/Shared/training-experiments
Code source configuration
code_source:
type: snapshot
snapshot:
root_path: /home/username/repo # REQUIRED — local path to repo or directory
git: # Optional (git repos only) — pin to a branch or commit
branch: main # Uses the branch's local HEAD
# commit: abc1234567 # Mutually exclusive with 'branch'
remote: false # Optional; remote fetching is not supported
include_paths: # Optional — filter included paths
- src/
- configs/
Field constraints:
git.branchandgit.commitare mutually exclusive: specify exactly one within thegit:block.- Omit
git.remoteor set it tofalse. The CLI does not fetch revisions from a remote. - If you omit the
git:block, the working tree is packaged while honoring Git ignore rules, including uncommitted changes in the selected files.
Custom parameters
Passed to the workload via HYPERPARAMETERS_PATH:
parameters:
model:
name: 'gpt2'
hidden_size: 768
training:
batch_size: 32
MLflow run name
Use 1 to 100 ASCII letters, digits, hyphens, or underscores. If omitted, the run name defaults to experiment_name.
mlflow_run_name: 'experiment-001-baseline'
MLflow artifact location
Set mlflow_artifact_location to store artifacts for an MLflow experiment in a custom root location. If you omit this field, a new experiment uses the default DBFS location, such as dbfs:/databricks/mlflow-tracking/<experiment-id>/....
mlflow_artifact_location: /Volumes/main/default/mlflow-artifacts/my-training
If DBFS access is restricted or you prefer Unity Catalog, specify either a /Volumes/<catalog>/<schema>/<volume>/... path or the equivalent dbfs:/Volumes/<catalog>/<schema>/<volume>/... URI. The Databricks CLI converts a /Volumes path to the dbfs: URI that MLflow uses.
Use a location unique to each experiment. An MLflow experiment's artifact location is fixed when the experiment is created. If experiment_name identifies an existing experiment, mlflow_artifact_location must match its artifact location or be omitted. To use a different location, specify a new experiment name.
Path resolution
Relative code_source.snapshot.root_path values resolve from the workload YAML file's directory. include_paths entries resolve from root_path. Paths inside command refer to files on the remote runtime, not your local machine. Use $CODE_SOURCE_PATH to reference uploaded code.
Folder structure:
/home/username/my-project/
├── train.yaml
└── scripts/
└── train.py
YAML configuration:
experiment_name: my-training
environment:
dependencies:
- torch
- transformers
compute:
num_accelerators: 8
accelerator_type: GPU_8xH100
code_source:
type: snapshot
snapshot:
root_path: . # Relative to train.yaml
git:
branch: main
command: torchrun --nproc_per_node=8 $CODE_SOURCE_PATH/scripts/train.py